Target pre-binding method and target pre-binding device based on cross-target key point

Through deep neural networks to identify and match the central prediction targets and potential binding prediction targets in the target image, the accuracy of binding relationship recognition is solved when multiple targets overlap is high, and a more stable and efficient target binding relationship prediction is achieved.

CN119942425APending Publication Date: 2025-05-06CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411983532.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When multiple targets overlap high, it is difficult for the prior art to accurately identify the binding relationship between targets, resulting in a high error binding rate.

Method used

The target binding method based on deep neural network is adopted to extract features in the target image, identify the central prediction target and potential binding prediction target, and use the bounding box matching calculation and matching algorithm to determine the binding relationship between the target.

Benefits of technology

In complex scenarios, the binding relationship between targets can be more stable and accurate, the error binding rate can be reduced, and the effect of intelligent algorithms can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942425A_ABST
    Figure CN119942425A_ABST
Patent Text Reader

Abstract

The invention provides a target pre-binding method and a target pre-binding device based on a cross-target key point, and the method comprises the steps: obtaining a target image which is an image obtained after the preprocessing of an original image; extracting features in the target image based on a deep neural network to obtain a plurality of prediction targets and a plurality of pieces of prediction information corresponding to the prediction targets; screening all the prediction information corresponding to each prediction target to obtain target prediction information corresponding to the prediction target, and enabling one prediction target to correspond to one piece of target prediction information; performing bounding box matching calculation according to the target prediction information of each prediction target to obtain a corresponding matching cost; and matching by using a matching algorithm according to the matching cost to obtain a binding relationship between the central prediction target and the potential binding prediction target. According to the method and the device, the problem that the binding relationship between the targets cannot be accurately identified under the condition that the overlapping degree of the multiple targets is relatively high in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target binding, and in particular to a target pre-binding method based on cross-target key points, a target pre-binding device, a computer-readable storage medium and a computer program product. Background Art

[0002] With the rapid development of artificial intelligence and computer vision technology, intelligent analysis and intelligent monitoring of images and videos have shined in many fields such as security, transportation, and urban governance. In a series of algorithm tasks such as passenger flow statistics, face capture, vehicle capture, illegal parking recognition, and motor vehicle driver violations, in addition to the detection and tracking of related targets, it is also necessary to identify a variety of binding relationships between targets, such as the binding between the face, head, and human body target frames of the same person, the binding between the vehicle and its license plate, and the binding between the vehicle and the driver and passengers.

[0003] Usually, target detection and target binding are divided into two steps, one before and one after. That is, after the target detection model identifies the type and bounding box of the target, the target post-binding algorithm uses the intersection-over-union ratio between the bounding boxes as the matching cost, and uses traditional matching algorithms such as the Hungarian algorithm to predict the binding relationship between the targets. This method is usually effective and concise when the targets on the image are sparse and there is no serious overlap between multiple targets. However, in complex scenes, when there is a serious overlap between multiple targets (such as binding the human body frame and the head frame in an image of an adult holding a child), post-binding that only considers the intersection-over-union ratio of the target frame will often fail, resulting in a high misbinding rate. Summary of the invention

[0004] The main purpose of the present application is to provide a target pre-binding method, a target pre-binding device, a computer-readable storage medium and a computer program product based on cross-target key points, so as to at least solve the problem in the prior art that the binding relationship between targets cannot be accurately identified when the overlap degree of multiple targets is high.

[0005] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a target pre-binding method based on cross-target key points is provided, including: obtaining a target image, wherein the target image is an image after preprocessing the original image; extracting features in the target image based on a deep neural network, obtaining multiple prediction targets and multiple prediction information corresponding to the prediction targets, wherein the prediction targets are divided into a central prediction target and a potential binding prediction target, wherein the central prediction target is a subject to be predicted on the target image, and the potential binding prediction target is a target that has a potential interactive relationship with the central prediction target, and each of the prediction information includes the confidence, category, key point coordinates, key point flag and bounding box corresponding to the prediction target, and the key point flag is used to predict the Whether the predicted target has key points, the bounding box is used to select the position of the predicted target in the target image, and the key point represents the key position of the predicted target on the bounding box; all the prediction information corresponding to each of the predicted targets is screened to obtain the target prediction information corresponding to the predicted target, and one predicted target corresponds to one target prediction information; according to the target prediction information of each of the predicted targets, a bounding box matching calculation is performed to obtain the corresponding matching cost, and the matching cost is the Euclidean distance between the bounding box of the central predicted target and the bounding box of the potential binding predicted target; matching is performed using a matching algorithm according to the matching cost to obtain the binding relationship between the central predicted target and the potential binding predicted target.

[0006] Optionally, obtaining a target image includes: obtaining an original image; converting the original image into a target image format to obtain a first image, wherein the target image format is at least in RGB format; adjusting the size of the first image to a target size and performing edge filling to obtain a second image; converting the second image into a tensor form and performing normalization to obtain the target image.

[0007] Optionally, before extracting features in the target image based on a deep neural network to obtain multiple predicted targets and multiple prediction information corresponding to the predicted targets, the method further includes: obtaining an initial sample target detection data set, and marking a sample binding relationship for each sample in the initial sample target detection data set, the initial sample target detection data set includes multiple sample targets, each of which includes a sample category, a sample bounding box, and a sample confidence; generating corresponding sample key points and sample key point flags of the sample bounding box according to each of the sample binding relationships, and updating the initial sample target detection data set according to the sample key points and the sample key point flags to obtain a sample target detection data set; a training step, The sample targets in the sample target detection data set are input into an initial deep neural network, and the initial deep neural network is iteratively trained using a target loss function to obtain a current deep neural network, wherein the initial deep neural network is an untrained deep neural network, and the target loss function at least includes a key point landmark classification loss function and a key point coordinate regression loss function; a testing step, using a test set to test the current deep neural network to obtain test indicators, wherein the test indicators at least include accuracy and false alarm rate; a repetition step, when the test indicators do not meet the set accuracy, adding sample targets or adjusting hyperparameters, and repeating the training step and the testing step in sequence at least once until the test indicators meet the set accuracy.

[0008] Optionally, all the prediction information corresponding to each of the prediction targets are screened to obtain target prediction information corresponding to the prediction target, including: mapping the confidence in each of the prediction information through a Sigmoid activation function to obtain a corresponding target confidence, wherein the target confidence is a value between 0 and 1; normalizing the category in each of the prediction information through a normalized exponential function to obtain a corresponding target category vector and a target category, wherein the elements in the target category vector are all values ​​between 0 and 1 and the sum of all the elements is equal to 1; retaining the potential binding prediction target whose target confidence is greater than or equal to a confidence threshold to obtain a binding prediction target; converting the coordinates of the bounding box in all the prediction information corresponding to the binding prediction target and the central prediction target from local coordinates to global coordinates on the target image; and using a non-maximum suppression algorithm to retain the prediction information with the highest target confidence for each of the prediction targets to obtain target prediction information corresponding to each of the prediction targets.

[0009] Optionally, after screening all the prediction information corresponding to each of the prediction targets to obtain the target prediction information corresponding to the prediction target, the method further includes: when there is only one binding prediction target and the key point flag of the central prediction target is 0, determining that the central prediction target has no binding relationship in the target image, the binding prediction target is the potential binding prediction target whose target confidence is greater than or equal to a confidence threshold, and the target confidence is the confidence after mapping processing of the confidence in the prediction information.

[0010] Optionally, a bounding box matching calculation is performed according to the target prediction information of each of the prediction targets to obtain a corresponding matching cost, including: in the case where there are multiple binding prediction targets, the Euclidean distance between the key point of the bounding box of each of the binding prediction targets and the center point of the bounding box of the central prediction target is calculated to obtain the corresponding matching cost, the binding prediction target is the potential binding prediction target whose target confidence is greater than or equal to a confidence threshold, the target confidence is the confidence after mapping the confidence in the prediction information, and the center point is the center position of the bounding box; in the case where there is only one binding prediction target and the key point flag bit of the central prediction target is 1, the average of the first Euclidean distance and the second Euclidean distance is calculated to obtain the matching cost, the first Euclidean distance is the Euclidean distance between the key point of the bounding box of the binding prediction target and the center point of the bounding box of the central prediction target, and the second Euclidean distance is the Euclidean distance between the key point of the bounding box of the central prediction target and the center point of the bounding box of the binding prediction target.

[0011] Optionally, the matching algorithm includes at least a greedy algorithm, a depth-first search algorithm and a Hungarian algorithm, and matching is performed using the matching algorithm according to the matching cost to obtain a binding relationship between the central prediction target and the potential binding prediction target, including: using a first matching algorithm for matching to obtain a first binding relationship, and the first matching algorithm is at least one of the following: a greedy algorithm, a depth-first search algorithm and a Hungarian algorithm; using a second matching algorithm for matching to obtain a second binding relationship, and the second matching algorithm is the matching algorithm other than the first matching algorithm; when the first binding relationship is the same as the second binding relationship, determining the first binding relationship as the binding relationship between the central prediction target and the potential binding prediction target; when the first binding relationship is different from the second binding relationship, using a third matching algorithm for matching to obtain a third binding relationship, and the third matching algorithm is the matching algorithm other than the first matching algorithm and the second matching algorithm; when the third binding relationship is the same as the first binding relationship or the third binding relationship is the same as the second binding relationship, determining the third binding relationship as the binding relationship between the central prediction target and the potential binding prediction target.

[0012] According to another aspect of the present application, a target pre-binding device based on cross-target key points is provided, and the device includes: a first acquisition unit, used to acquire a target image, and the target image is an image after preprocessing the original image; an extraction unit, used to extract features in the target image based on a deep neural network, and obtain multiple prediction targets and multiple prediction information corresponding to the prediction targets, the prediction targets are divided into a central prediction target and a potential binding prediction target, the central prediction target is the subject to be predicted on the target image, and the potential binding prediction target is a target that has a potential interactive relationship with the central prediction target, each of the prediction information includes the confidence, category, key point coordinates, key point flag and bounding box corresponding to the prediction target, and the key point flag is used to predict that the prediction target is Whether there are key points, the bounding box is used to frame the position of the predicted target in the target image, and the key point represents the key position of the predicted target on the bounding box; a target detection unit is used to screen all the prediction information corresponding to each of the predicted targets to obtain the target prediction information corresponding to the predicted target, one predicted target corresponds to one target prediction information; a calculation unit is used to perform bounding box matching calculation according to the target prediction information of each predicted target to obtain the corresponding matching cost, and the matching cost is the Euclidean distance between the bounding box of the central predicted target and the bounding box of the potential bound prediction target; a matching unit is used to perform matching using a matching algorithm according to the matching cost to obtain the binding relationship between the central predicted target and the potential bound prediction target.

[0013] According to another aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute any one of the methods described.

[0014] According to another aspect of the present application, a computer program product is provided, comprising computer instructions, wherein when the computer instructions are executed by a processor, any one of the methods described above is implemented.

[0015] Applying the technical solution of the present application, in a target pre-binding method based on cross-target key points, first, a target image is obtained, and the target image is an image after preprocessing the original image; then, based on a deep neural network, features in the target image are extracted to obtain multiple prediction targets and multiple prediction information corresponding to the prediction targets. The prediction targets are divided into central prediction targets and potential binding prediction targets. The central prediction target is the subject to be predicted on the target image, and the potential binding prediction target is a target that has a potential interactive relationship with the central prediction target. Each of the prediction information includes the confidence, category, key point coordinates, key point flag and bounding box corresponding to the prediction target. The key point flag is used to predict whether the prediction target is related Key point, the above-mentioned bounding box is used to select the position of the above-mentioned predicted target in the above-mentioned target image, and the above-mentioned key point represents the key position of the above-mentioned predicted target on the above-mentioned bounding box; then, all the above-mentioned prediction information corresponding to each of the above-mentioned predicted targets is screened to obtain the target prediction information corresponding to the above-mentioned predicted target, and one above-mentioned predicted target corresponds to one above-mentioned target prediction information; then, the bounding box matching calculation is performed according to the above-mentioned target prediction information of each of the above-mentioned predicted targets to obtain the corresponding matching cost, and the above-mentioned matching cost is the Euclidean distance between the bounding box of the above-mentioned central predicted target and the bounding box of the above-mentioned potential binding prediction target; finally, matching is performed using the matching algorithm according to the above-mentioned matching cost to obtain the binding relationship between the above-mentioned central predicted target and the above-mentioned potential binding prediction target. The key point of this application is cross-target, that is, it can predict the position of other targets (i.e., potential binding prediction targets) that have potential interactive relationships with the central prediction target, so that while identifying the categories and positions of the central prediction target and the potential binding prediction target, the binding relationship between the targets can be predicted. This method breaks through the existing limitation that target binding can only be performed through post-processing, so that the binding relationships such as subordination and interaction between targets can be intelligently predicted through deep learning models. Especially when the overlap of multiple targets is high, the binding relationship can be identified more stably and accurately, thereby improving the effect of each intelligent algorithm in complex scenarios. This application solves the problem in the prior art that the binding relationship between targets cannot be accurately identified when the overlap of multiple targets is high. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A hardware structure block diagram of a mobile terminal for executing a target pre-binding method based on cross-target key points provided in an embodiment of the present application is shown;

[0017] Figure 2 A flow chart of a target front binding method based on cross-target key points provided according to an embodiment of the present application is shown;

[0018] Figure 3A structural diagram of a target pre-binding system based on cross-target key points provided according to an embodiment of the present application is shown;

[0019] Figure 4 A schematic diagram showing target overlap resulting in post-binding misbinding using the traditional method;

[0020] Figure 5 A schematic diagram showing the effect of a target front binding based on cross-target key points provided according to an embodiment of the present application is shown;

[0021] Figure 6 A structural block diagram of a target pre-binding device based on cross-target key points provided according to an embodiment of the present application is shown.

[0022] The above drawings include the following reference numerals:

[0023] 102, processor; 104, memory; 106, transmission device; 108, input and output devices. DETAILED DESCRIPTION

[0024] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0025] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0027] For the convenience of description, some nouns or terms involved in the embodiments of the present application are explained below:

[0028] Object detection: refers to identifying the category and location of the desired object in an image, which is one of the core tasks of computer vision.

[0029] Target binding: refers to matching multiple targets detected in an image one-to-one or one-to-many according to a certain relationship, such as matching the body frame of the same person with the head frame, matching the vehicle detection frame with its license plate detection frame, etc. Usually, after the target detection model identifies the category and location of the target, a post-processing method, such as the Hungarian matching algorithm, is used to predict the matching relationship between targets. This method is called "post-target binding".

[0030] Pre-target binding: refers to using a deep learning model to identify the binding relationship between targets while identifying the target category and location. That is, a detection model can detect targets and bind targets. It can effectively solve the problem of high misbinding rate in post-target binding under certain circumstances.

[0031] As introduced in the background technology, in the prior art, in complex scenarios, when there is a serious overlap between multiple targets, the post-binding that only considers the target frame intersection and union ratio often fails, resulting in a high misbinding rate. In order to solve the problem in the prior art that the binding relationship between targets cannot be accurately identified when the overlap between multiple targets is high, the embodiments of the present application provide a target pre-binding method based on cross-target key points, a target pre-binding device, a computer-readable storage medium and a computer program product.

[0032] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0033] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 1 is a hardware structure block diagram of a mobile terminal according to an embodiment of the present invention, based on a target pre-binding method across target key points. Figure 1 As shown, the mobile terminal may include one or more ( Figure 1 Only one is shown in the figure) a processor 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data, wherein the mobile terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the mobile terminal. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations shown.

[0034] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the target pre-binding method based on cross-target key points in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, the above method is realized. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories can be connected to the mobile terminal via a network. Examples of the above-mentioned network include but are not limited to the Internet, corporate intranets, local area networks, mobile communication networks and combinations thereof. The transmission device 106 is used to receive or send data via a network. The above-mentioned network specific examples may include a wireless network provided by a communication provider of the mobile terminal. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0035] In this embodiment, a target pre-binding method based on cross-target key points that runs on a mobile terminal, a computer terminal or a similar computing device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0036] Figure 2 is a flow chart of a target front binding method based on cross-target key points according to an embodiment of the present application. Figure 2 As shown, the method comprises the following steps:

[0037] Step S201, obtaining a target image, wherein the target image is an image obtained by preprocessing an original image.

[0038] Specifically, the original images may come from a variety of different scenes and devices, including different resolutions, color formats, lighting conditions, etc. The goal of preprocessing is to convert these images into a unified format that can be processed by deep neural networks. Preprocessing may include reading the original image, converting the color space (such as from BGR to RGB), adjusting the image size, performing edge padding, normalization, and other operations to obtain the target image to ensure the consistency of the model input data.

[0039] Step S202, based on the deep neural network, extract the features in the above-mentioned target image to obtain multiple prediction targets and multiple prediction information corresponding to the above-mentioned prediction targets, the above-mentioned prediction targets are divided into central prediction targets and potential binding prediction targets, the above-mentioned central prediction target is the subject to be predicted on the above-mentioned target image, the above-mentioned potential binding prediction target is the target that has a potential interactive relationship with the above-mentioned central prediction target, each of the above-mentioned prediction information includes the confidence, category, key point coordinates, key point flag and bounding box corresponding to the above-mentioned prediction target, the above-mentioned key point flag is used to predict whether the above-mentioned prediction target has key points, the above-mentioned bounding box is used to frame the position of the above-mentioned prediction target in the above-mentioned target image, and the above-mentioned key points represent the key positions of the above-mentioned prediction targets on the above-mentioned bounding box.

[0040] Specifically, Figure 3 As shown, the target pre-binding system of the present invention includes a pre-processing module, an image analysis module and a post-processing module. Through the image analysis module, based on the pre-processed target image, the deep neural network performs feature extraction and identifies multiple targets in the image. These targets are classified into central prediction targets and potential binding prediction targets. The prediction information includes confidence (probability of target existence), category (target type), key point coordinates, key point flags and bounding boxes. The deep neural network learns the feature representation of the target through components such as convolutional layers, fully connected layers, and regression layers, and predicts the location, category and key point information of the target based on these features. Through the deep neural network, the model can automatically learn the visual features of the target without manually designing complex feature extraction algorithms. The confidence and category in the prediction information are used to determine the existence and type of the target, while the key point coordinates and flags are used for subsequent binding relationship prediction. The prediction of the bounding box ensures the precise positioning of the target and is the basis of the target detection task.

[0041] Step S203, screening all the above prediction information corresponding to each of the above prediction targets to obtain target prediction information corresponding to the above prediction targets, where one prediction target corresponds to one target prediction information.

[0042] Specifically, after obtaining multiple prediction targets and their prediction information, it is necessary to filter this information and only retain the target prediction information with high confidence. The screening process usually includes setting a confidence threshold, retaining only the prediction information with confidence higher than the threshold, and performing non-maximum suppression on multiple prediction results of similar targets to eliminate overlapping prediction boxes. Each prediction target corresponds to one target prediction information. By setting the confidence threshold, low-quality detection results can be filtered out, improving the accuracy of the final binding relationship prediction.

[0043] Step S204 , performing bounding box matching calculation according to the target prediction information of each of the predicted targets, and obtaining a corresponding matching cost, wherein the matching cost is the Euclidean distance between the bounding box of the central predicted target and the bounding box of the potential bounding predicted target.

[0044] Specifically, for the central prediction target and the potential binding prediction target, the Euclidean distance between the bounding box of the central prediction target and the bounding box of the potential binding prediction target is calculated as the matching cost. In the case of a one-to-one binding relationship, the matching cost is calculated as the average Euclidean distance between the center points of the two bounding boxes; in the case of a one-to-many binding relationship, the matching cost is the Euclidean distance from the center point of the central prediction target to the key points of other potential binding prediction targets. The Euclidean distance can quantify the spatial relationship between targets. The smaller the distance, the closer the two targets are, and the stronger the potential binding relationship. Through the matching cost, the algorithm can intelligently determine which potential binding prediction targets have a binding relationship with the central prediction target, thereby more accurately predicting the binding relationship and reducing the occurrence of misbinding and missed binding.

[0045] Step S205 , performing matching using a matching algorithm according to the matching cost, and obtaining a binding relationship between the central prediction target and the potential binding prediction target.

[0046] Specifically, according to the calculated matching cost, a matching algorithm (such as a greedy algorithm, a depth-first search algorithm, or a Hungarian algorithm, etc.) is used to match and obtain the binding relationship between the central prediction target and the potential binding prediction target. The goal of the matching algorithm is to find a set of matching relationships so that the total cost of all matches is minimized or other optimization goals are met. The matching algorithm is a decision-making process for binding relationship prediction, which determines the optimal binding relationship based on the matching cost. Through the intelligent algorithm, the target pair with the minimum matching cost can be automatically found, thereby achieving efficient and accurate binding relationship prediction.

[0047] It should be noted that, in the face of overlapping targets, the intersection-and-union ratio between target frames may not necessarily reflect the binding relationship between targets (for example, the intersection-and-union ratio of the target frame of an electric vehicle and its rider may be smaller than the intersection-and-union ratio of the electric vehicle and the pedestrian next to it), resulting in a high misbinding rate in the traditional target post-binding algorithm, which further affects the effects of various intelligent algorithms, such as Figure 4 As shown, Figure 4 A schematic diagram shows a mis-binding caused by overlapping targets bound by traditional methods. In the figure, A1 represents the key points of the head bounding box of predicted target A, A2 represents the key points of the human body bounding box of predicted target A, B1 represents the key points of the head bounding box of predicted target B, and B2 represents the key points of the human body bounding box of predicted target B. Mis-binding problems occur when binding by traditional methods. However, the method of the present invention can effectively identify binding relationships in which target frames do not overlap. For a binding relationship such as a pet dog and the owner who walks the dog, the two target frames may not overlap at all, causing the post-binding algorithm based on intersection-union ratio to be completely invalid, while the pre-target binding algorithm based on cross-target key points is basically unaffected. Figure 5 As shown, Figure 5 A schematic diagram of the effect of a cross-target key point pre-binding provided according to an embodiment of the present application is shown, in which A1 represents the key points of the head bounding box of the predicted target A, A2 represents the key points of the human body bounding box of the predicted target A, B1 represents the key points of the head bounding box of the predicted target B, and B2 represents the key points of the human body bounding box of the predicted target B. No mis-binding problem occurs when binding using the method of the present invention. The present invention can be applied to the recognition of various binding relationships such as face and human body, head and human body, vehicle and license plate, vehicle and driver, etc., effectively solving the mis-binding problem of the post-target binding algorithm in complex scenes, thereby improving the effect of the intelligent recognition algorithm in complex scenes.

[0048] In the present embodiment, in the target pre-binding method based on cross-target key points, first, a target image is obtained, and the target image is an image after preprocessing the original image; then, based on the deep neural network, features in the target image are extracted to obtain multiple prediction targets and multiple prediction information corresponding to the prediction targets. The prediction targets are divided into central prediction targets and potential binding prediction targets. The central prediction target is the subject to be predicted on the target image, and the potential binding prediction target is a target that has a potential interactive relationship with the central prediction target. Each of the prediction information includes the confidence, category, key point coordinates, key point flag and bounding box corresponding to the prediction target. The key point flag is used to predict whether the prediction target has a key point. , the above-mentioned bounding box is used to select the position of the above-mentioned predicted target in the above-mentioned target image, and the above-mentioned key point represents the key position of the above-mentioned predicted target on the above-mentioned bounding box; then, all the above-mentioned prediction information corresponding to each of the above-mentioned predicted targets is screened to obtain the target prediction information corresponding to the above-mentioned predicted target, and one above-mentioned predicted target corresponds to one above-mentioned target prediction information; then, the bounding box matching calculation is performed according to the above-mentioned target prediction information of each of the above-mentioned predicted targets to obtain the corresponding matching cost, and the above-mentioned matching cost is the Euclidean distance between the bounding box of the above-mentioned central predicted target and the bounding box of the above-mentioned potential binding prediction target; finally, matching is performed using the matching algorithm according to the above-mentioned matching cost to obtain the binding relationship between the above-mentioned central predicted target and the above-mentioned potential binding prediction target. The key point of this application is cross-target, that is, it can predict the position of other targets (i.e., potential binding prediction targets) that have potential interactive relationships with the central prediction target, so that while identifying the categories and positions of the central prediction target and the potential binding prediction target, the binding relationship between the targets can be predicted. This method breaks through the existing limitation that target binding can only be performed through post-processing, so that the binding relationships such as subordination and interaction between targets can be intelligently predicted through a deep learning model. Especially when the overlap of multiple targets is high, the binding relationship can be identified more stably and accurately, thereby improving the effect of each intelligent algorithm in complex scenarios. This application solves the problem in the prior art that the binding relationship between targets cannot be accurately identified when the overlap of multiple targets is high.

[0049] In order to enable those skilled in the art to more clearly understand the technical solution of the present application, the implementation process of the target pre-binding method based on cross-target key points of the present application will be described in detail below in combination with specific embodiments.

[0050] In order to improve the quality of image data, in an optional implementation manner, the above step S201 includes:

[0051] Step S2011, obtaining an original image;

[0052] Step S2012, converting the original image into a target image format to obtain a first image, wherein the target image format is at least an RGB format;

[0053] Step S2013, adjusting the size of the first image to a target size and performing edge filling to obtain a second image;

[0054] Step S2014, converting the second image into a tensor form and performing normalization processing to obtain the target image.

[0055] In the above embodiment, if Figure 3 As shown, the target front binding system of the present invention includes a pre-processing module, an image analysis module and a post-processing module. The function of the pre-processing module is to convert images with different resolutions, aspect ratios and compression protocols into a unified format required by the image analysis module. First, read the image file, which can be a static image directly read through the file system, or intercept a frame from the video stream. In addition, the original image may come from a variety of different devices, such as surveillance cameras, mobile phone cameras, etc., so their resolutions, aspect ratios and compression methods may be different. Convert the original image to RGB format or other formats to obtain the first image, because the RGB format is based on the color space of the red, green and blue channels, and the deep learning model usually trains RGB images, so it is necessary to convert the image data into a format that the model can understand. Adjust the image size to the input size required by the model by scaling, cropping or stretching, etc., usually this target size is a fixed size defined during model training. Then, fill the edges of the adjusted image to obtain the second image to avoid image information loss caused by image scaling. The image size is adjusted to meet the input requirements of the model, so that the model is based on the same size input when processing all images, thereby simplifying the calculation and improving the processing speed. Edge padding helps maintain the integrity of the image and avoids the loss of key information due to resizing, especially when the target is at the edge of the image. The second image is converted into a tensor, which is a common data structure in deep learning frameworks that can perform mathematical operations efficiently. Next, the pixel values ​​in the tensor are normalized, for example, the RGB pixel values ​​are scaled from the range of 0 to 255 to the range of 0 to 1, which helps with model training and prediction because the normalized data can speed up model convergence and reduce numerical instability. It ensures that the image data has been standardized and optimized before entering the deep learning model, so that the model can perform target detection and binding relationship prediction more effectively. This not only improves the computational efficiency of the model, but also enhances the generalization ability of the model under different image conditions, which is especially important for processing image data in complex scenes.

[0056] In order to improve the accuracy and stability of target detection and analysis, in an optional implementation, before the above step S202, the method further includes:

[0057] Step S301, obtaining an initial sample target detection data set, and marking a sample binding relationship for each sample in the initial sample target detection data set, wherein the initial sample target detection data set includes a plurality of sample targets, and each of the sample targets includes a sample category, a sample bounding box, and a sample confidence;

[0058] Step S302, generating the sample key points and sample key point flags of the corresponding sample bounding boxes according to the sample binding relationships, and updating the initial sample target detection data set according to the sample key points and the sample key point flags to obtain a sample target detection data set;

[0059] Step S303, a training step, inputting the sample targets in the sample target detection data set into an initial deep neural network, and iteratively training the initial deep neural network using a target loss function to obtain a current deep neural network, wherein the initial deep neural network is an untrained deep neural network, and the target loss function at least includes a key point landmark classification loss function and a key point coordinate regression loss function;

[0060] Step S304, a testing step, using a test set to test the current deep neural network to obtain a test index, wherein the test index at least includes an accuracy rate and a false alarm rate;

[0061] Step S305, repeating the steps, if the above test indicators do not meet the set accuracy, increase the sample target or adjust the hyperparameters, and repeat the above training steps and the above test steps in sequence at least once until the above test indicators meet the set accuracy.

[0062] In the above embodiments, usually, the target detection data set does not have the annotation of the binding relationship, and the method proposed in this patent requires the model to predict the target category, bounding box and binding relationship at the same time, so the target detection data set (including training and test sets) is required. First, an initial target detection data set is obtained from an existing database or acquisition system. This data set contains multiple image samples, and each sample has the target category, bounding box and confidence information marked. Then, for the targets in each sample, the binding relationship between them is manually or using auxiliary software. For example, in an image containing multiple human bodies and human heads, it is necessary to mark the binding relationship between each head frame and the corresponding human body frame. Basic data for learning binding relationships is provided, so that the model can learn the association between target detection and binding relationship prediction, so that accurate binding prediction can be performed on unknown data. Based on the marked binding relationship, a key point is generated for each target frame, and this key point points to the center of other targets bound to it. At the same time, a key point flag is generated for each target frame to indicate whether the target frame needs to be bound. Then, these key point coordinates and flag information are added to the sample data set to update the original data set. The generation of keypoints and keypoint landmarks is a key data preparation process for the pre-training binding algorithm model. Keypoints provide a direct indication of the binding relationship between objects, while landmarks help the model understand which object boxes need to be bound. The updated sample object detection dataset is input into the initial deep neural network, which can be YOLO, Faster R-CNN, or any other deep learning architecture suitable for object detection. Then, the target loss function is defined, which includes at least the keypoint landmark classification loss function (such as BCELoss, BinaryCrossEntropyLoss, Binary Cross Entropy Loss) and the keypoint coordinate regression loss function (such as L2Loss, MeanSquared Error Loss, Mean Squared Error Loss). In each iteration, the model predicts the object bounding box, category, keypoint coordinates, and keypoint landmarks compared with the true value of the sample, calculates the loss, and then uses the optimizer (such as Adam) to update the network parameters to minimize the loss function. Through iterative training, the model gradually learns to predict the category and location of the object from the input image, and can also predict the keypoint coordinates and landmarks, and then predict the binding relationship between two object boxes. The definition of the loss function ensures that the model can optimize the performance of both object detection and binding relationship prediction, thereby improving the accuracy and robustness of the overall model. After the model is trained, it is evaluated using an independent test set that was not involved in the training. The test set contains image samples that are also labeled with object categories, bounding boxes, and binding relationships.The model predicts the images in the test set, outputs the predicted target box, category, key point coordinates and key point flags, and then compares them with the true values ​​to calculate the predicted accuracy and false alarm rate and other indicators. Indicators such as accuracy and false alarm rate can reflect the performance of the model in target detection and binding relationship prediction, and help determine whether the model has reached the expected accuracy standard. If the test indicators do not meet the preset accuracy standards, it is necessary to further optimize the model by adding more sample data or adjusting the model's hyperparameters (such as learning rate, regularization term). Specifically, samples containing complex binding relationships can be added to improve the robustness of the model, or the learning rate can be adjusted, the number of training rounds can be increased, etc., to improve the training effect of the model. Then, the above training steps and the above test steps are repeated in sequence until the test indicators of the model meet the set accuracy, and finally a deep neural network model that performs well in both target detection and binding relationship prediction is obtained. Through this process, the model can learn and predict complex binding relationships, thereby realizing a more accurate and efficient front binding algorithm in target detection tasks, especially in complex scenes with high target overlap, which can significantly improve the accuracy and stability of target detection and analysis.

[0063] In order to improve the accuracy of the detection result, in an optional implementation manner, the above step S203 includes:

[0064] Step S2031, mapping the confidences in each of the above prediction information through a Sigmoid activation function to obtain a corresponding target confidence, where the target confidence is a value between 0 and 1;

[0065] Step S2032, normalizing the categories in each of the above prediction information by using a normalized exponential function to obtain a corresponding target category vector and a target category, wherein the elements in the above target category vector are all values ​​between 0 and 1 and the sum of all the above elements is equal to 1;

[0066] Step S2033, retaining the potential binding prediction targets whose target confidence is greater than or equal to the confidence threshold, to obtain binding prediction targets;

[0067] Step S2034, converting the coordinates of the bounding boxes in all the prediction information corresponding to the bounding prediction target and the central prediction target from local coordinates to global coordinates on the target image;

[0068] Step S2035, using a non-maximum suppression algorithm, retains the prediction information with the highest target confidence for each of the prediction targets, and obtains target prediction information corresponding to each of the prediction targets.

[0069] In the above embodiment, if Figure 4As shown in the figure, the main function of the post-processing module is to convert the output of the image analysis module into the required prediction values, such as confidence, category, bounding box, binding relationship, etc. The post-processing module is divided into a target detection post-processing module and a binding relationship post-processing module. The output of the deep neural network usually contains the original confidence value (usually ranging from negative infinity to positive infinity). These values ​​need to be converted through the Sigmoid activation function to map them between 0 and 1, so that the target confidence obtained can represent the model's probability estimate of the existence of a certain target in the image. The categories in each of the above prediction information are normalized by the normalized exponential function to obtain the corresponding target category vector and target category implementation process: The category information output by the model is usually a set of original category scores. By using the normalized exponential function (usually referring to the Softmax function), these scores are converted into probability vectors, where the sum of the probabilities of all categories is equal to 1. Based on this probability vector, the category with the highest probability can be selected as the final target category. Based on the obtained target confidence, a confidence threshold is set to retain only those potential binding targets with confidence higher than the threshold. Function and effect: Confidence threshold filtering is a key step to remove low-quality detection frames. By setting a reasonable threshold, detection results that the model is not sure about can be excluded, improving the accuracy and reliability of the detected target. The predictions of deep learning models are usually given in the form of local coordinates, that is, the coordinates relative to the grid cell to which the detection frame belongs. In order to accurately mark the position of the detection frame on the original image, it is necessary to convert these local coordinates into global coordinates, that is, the actual coordinate position on the image. Coordinate conversion ensures the correct position of the detection frame on the original image, which is crucial for the visualization of target detection and binding relationships and subsequent processing. The non-maximum suppression (NMS) algorithm is used to filter out those overlapping bounding boxes. Specifically, for each predicted target, the non-maximum suppression algorithm compares the confidence of all bounding boxes of the predicted target, retains the box with the highest confidence and low overlap with other boxes, and removes the rest. The non-maximum suppression algorithm solves the problem of multiple detections of the same target, ensuring that each target is detected only once and the detection result has the highest confidence. It realizes the effective transformation from model output to target detection and binding relationship prediction results. The conversion of confidence and category probability, confidence threshold filtering, coordinate transformation, and the application of the NMS algorithm together constitute a complete post-processing process, which aims to optimize the accuracy of the detection results, reduce redundancy, and improve the robustness of the model. This ensures that the model can not only accurately detect the target, but also effectively identify the binding relationship between the targets, especially in the case of dense targets and high overlap in the image, which can significantly improve the efficiency and accuracy of the overall detection and analysis.

[0070] In order to significantly reduce false positives and false negatives in binding relationship prediction, in an optional implementation, after the above step S203, the method further includes:

[0071] Step S401, when there is only one binding prediction target and the key point flag of the central prediction target is 0, it is determined that the central prediction target has no binding relationship in the target image, the binding prediction target is the potential binding prediction target whose target confidence is greater than or equal to the confidence threshold, and the target confidence is the confidence after mapping the confidence in the prediction information.

[0072] In the above embodiment, multiple potential binding prediction targets have been obtained, which are prediction targets whose target confidence is greater than or equal to the confidence threshold. These targets are candidate targets that the model believes exist in the current image and may have a binding relationship with other targets. In a one-to-one binding task, as long as there is a binding relationship, the key point flag must be 1. If the key point flag is 0, it means that the model has not predicted the binding key points related to the central target, indicating that the target may be independent in the current scene and does not require a binding relationship. Therefore, it is confirmed that in the current scene, the model only predicts one type of binding target. For example, in the binding scene of the human body and the human head, if the model only predicts the human head as a potential binding target, and the key point flag of the central prediction target is 0, then it can be determined that the central prediction target has no binding relationship in the target image. The false positives and false negatives of the binding relationship prediction are significantly reduced, and the overall accuracy of the binding relationship recognition is improved.

[0073] It should be noted that for a one-to-one target binding task, the key point is defined as follows: if bounding box A and bounding box B are a pair of binding relationships, then the key point of A is the center point of B, and the key point of B is the center point of A; for a one-to-many target binding task, the key point is defined as follows: if bounding box A and bounding box B, bounding box A and bounding box C both form a binding relationship, then A has no key point, and the key points of B and C are the center point of A.

[0074] In order to accurately determine the binding relationship between two targets, in an optional implementation, the above step S204 includes:

[0075] Step S2041, in the case where there are multiple binding prediction targets, the Euclidean distance between the key points of the bounding boxes of the above-mentioned binding prediction targets and the center point of the above-mentioned bounding box of the above-mentioned central prediction target is calculated to obtain the corresponding matching cost, the above-mentioned binding prediction target is the above-mentioned potential binding prediction target whose target confidence is greater than or equal to the confidence threshold, the above-mentioned target confidence is the confidence after mapping processing of the confidence in the above-mentioned prediction information, and the above-mentioned center point is the center position of the above-mentioned bounding box;

[0076] Step S2042, when there is only one of the above-mentioned binding prediction targets and the above-mentioned key point flag bit of the above-mentioned central prediction target is 1, calculate the average of the first Euclidean distance and the second Euclidean distance to obtain the above-mentioned matching cost, the above-mentioned first Euclidean distance is the Euclidean distance between the above-mentioned key point of the above-mentioned enclosing box of the above-mentioned binding prediction target and the above-mentioned center point of the above-mentioned enclosing box of the above-mentioned central prediction target, and the above-mentioned second Euclidean distance is the Euclidean distance between the above-mentioned key point of the above-mentioned enclosing box of the above-mentioned central prediction target and the above-mentioned center point of the above-mentioned enclosing box of the above-mentioned binding prediction target.

[0077] In the above embodiment, all high-confidence potential binding prediction targets (confidence greater than or equal to the threshold) are obtained. These prediction targets are candidate targets that the model believes exist in the current image and may be bound to the central target. There are multiple binding prediction targets, which means it is a one-to-many binding task. For example, the binding between the three bounding boxes of the face, head, and body of the same person. For the central prediction target and each binding prediction target, the Euclidean distance between the key point of the binding prediction target and the center point of the central prediction target is calculated. The key point here is the point predicted by the deep learning network model in the detection phase to indicate the potential binding relationship, and the center point is the geometric center of the target bounding box. The calculated Euclidean distance is used as the matching cost to quantify the strength of the binding relationship between the central prediction target and the binding prediction target. The smaller the Euclidean distance, the stronger the binding relationship between the binding prediction target and the central prediction target, and the lower the matching cost. When there is only one binding prediction target (for example, the head frame is used as the binding target of the human frame) and the key point flag of the central prediction target (such as the human frame) is 1, that is, the model predicts that the central target needs to be bound. First, the first Euclidean distance from the key point of the binding prediction target to the center point of the central prediction target is calculated. Then, the second Euclidean distance from the key point of the central prediction target to the center point of the binding prediction target is calculated. The average of the first Euclidean distance and the second Euclidean distance is used as the final matching cost. This step takes into account the bidirectionality of the binding relationship and ensures that the judgment of the binding relationship is more accurate and fair. By calculating the distance between the key point and the center point as the matching cost, the binding relationship between the two targets can be judged more accurately, avoiding the limitations of the traditional method based on intersection and union ratio in complex scenes with overlapping targets. The calculation method of the Euclidean distance is not affected by the size and shape of the target frame, making the algorithm more stable when processing targets of different sizes and shapes, and reducing the possibility of misjudgment. This key point-based matching cost calculation strategy not only enhances the performance of the algorithm in complex scenes, but also simplifies the matching process and improves the processing efficiency of the algorithm.

[0078] In order to significantly enhance the reliability of binding relationship prediction, in an optional implementation manner, the above step S205 includes:

[0079] Step S2051, using a first matching algorithm to perform matching to obtain a first binding relationship, wherein the first matching algorithm is at least one of the following: a greedy algorithm, a depth-first search algorithm, and a Hungarian algorithm;

[0080] Step S2052, using a second matching algorithm to perform matching to obtain a second binding relationship, wherein the second matching algorithm is the matching algorithm other than the first matching algorithm;

[0081] Step S2053, when the first binding relationship is the same as the second binding relationship, determining the first binding relationship as the binding relationship between the central prediction target and the potential binding prediction target;

[0082] Step S2054: when the first binding relationship is different from the second binding relationship, a third matching algorithm is used for matching to obtain a third binding relationship, where the third matching algorithm is a matching algorithm other than the first matching algorithm and the second matching algorithm;

[0083] Step S2055, when the third binding relationship is the same as the first binding relationship or the third binding relationship is the same as the second binding relationship, the third binding relationship is determined as the binding relationship between the central prediction target and the potential binding prediction target.

[0084] In the above embodiment, a first matching algorithm is used, usually one of a greedy algorithm, a depth-first search algorithm or a Hungarian algorithm, to make a preliminary judgment on the match between the central prediction target and the potential binding prediction target. These algorithms determine the best binding relationship based on the matching cost (such as Euclidean distance). The greedy algorithm may select the target pair with the smallest matching cost as the binding relationship; the depth-first search algorithm may explore all possible binding combinations and select the best one; the Hungarian algorithm is designed to solve the maximum weight matching in the bipartite graph, and it can find a set of matching relationships so that the total cost of all matches (i.e., the total distance) is minimized. The second matching algorithm is used for matching to obtain a second possible binding relationship. The second matching algorithm should be a choice other than the first matching algorithm. For example, if the first algorithm is a greedy algorithm, the second algorithm can be a depth-first search algorithm or a Hungarian algorithm. Compare the first binding relationship and the second binding relationship. If the binding relationships obtained by the two algorithms are the same, that is, they both point to the same set of target pairs, then it can be determined that this is the correct binding relationship between the central prediction target and the potential binding prediction target. If the binding relationships obtained by the first algorithm and the second algorithm are different, it means that there are different results caused by uncertainty or algorithm differences. At this time, a third matching algorithm is needed to perform further matching to obtain a third possible binding relationship. Compare the third binding relationship with the first binding relationship or the second binding relationship. If the result of the third algorithm is the same as the result of any previous algorithm, then this result is adopted as the binding relationship between the central prediction target and the potential binding prediction target. If the first binding relationship, the second binding relationship, and the third binding relationship are not the same, a prompt is issued that there may be no binding relationship, or a detection error problem has occurred. By adopting multiple matching algorithms and comparing the consistency of the results, the reliability of the binding relationship prediction can be significantly enhanced. If different algorithms produce consistent results, this indicates that the prediction results are robust and reduce the errors that may be caused by a single algorithm. Using different matching algorithms can cope with different types of binding relationships and scene complexities. For example, the greedy algorithm is suitable for simple situations, while the Hungarian algorithm is more suitable for complex matching problems. This diversity ensures that the algorithm can flexibly cope with various scenarios and improves the overall robustness. When the results of the initial two algorithms are inconsistent, the third algorithm is introduced to help resolve the conflict, reduce the uncertainty caused by algorithm selection, and improve the certainty of the binding relationship prediction. Although the use of multiple algorithms may increase the amount of calculation, in most cases, the algorithm in the first step can give reliable results and reduce the dependence on more complex algorithms. This strategy optimizes the use of computing resources as much as possible while ensuring accuracy.

[0085] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0086] The embodiment of the present application also provides a target pre-binding device based on cross-target key points. It should be noted that the target pre-binding device based on cross-target key points of the embodiment of the present application can be used to execute the target pre-binding method based on cross-target key points provided by the embodiment of the present application. The device is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware of predetermined functions. Although the device described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware is also possible and conceived.

[0087] The following introduces the target pre-binding device based on cross-target key points provided in an embodiment of the present application.

[0088] Figure 6 is a structural block diagram of a target pre-binding device based on cross-target key points according to an embodiment of the present application. Figure 6 As shown, the device comprises:

[0089] The first acquisition unit 10 is used to acquire a target image, where the target image is an image obtained by preprocessing the original image.

[0090] Specifically, the original images may come from a variety of different scenes and devices, including different resolutions, color formats, lighting conditions, etc. The goal of preprocessing is to convert these images into a unified format that can be processed by deep neural networks. Preprocessing may include reading the original image, converting the color space (such as from BGR to RGB), adjusting the image size, performing edge padding, normalization, and other operations to obtain the target image to ensure the consistency of the model input data.

[0091] The extraction unit 20 is used to extract features in the above-mentioned target image based on a deep neural network, and obtain multiple prediction targets and multiple prediction information corresponding to the above-mentioned prediction targets. The above-mentioned prediction targets are divided into a central prediction target and a potential binding prediction target. The above-mentioned central prediction target is the subject to be predicted on the above-mentioned target image, and the above-mentioned potential binding prediction target is a target that has a potential interactive relationship with the above-mentioned central prediction target. Each of the above-mentioned prediction information includes the confidence, category, key point coordinates, key point flag and bounding box corresponding to the above-mentioned prediction target. The above-mentioned key point flag is used to predict whether the above-mentioned prediction target has a key point. The above-mentioned bounding box is used to select the position of the above-mentioned prediction target in the above-mentioned target image, and the above-mentioned key point represents the key position of the above-mentioned prediction target on the above-mentioned bounding box.

[0092] Specifically, Figure 3 As shown, the target pre-binding system of the present invention includes a pre-processing module, an image analysis module and a post-processing module. Through the image analysis module, based on the pre-processed target image, the deep neural network performs feature extraction and identifies multiple targets in the image. These targets are classified into central prediction targets and potential binding prediction targets. The prediction information includes confidence (probability of target existence), category (target type), key point coordinates, key point flags and bounding boxes. The deep neural network learns the feature representation of the target through components such as convolutional layers, fully connected layers, and regression layers, and predicts the location, category and key point information of the target based on these features. Through the deep neural network, the model can automatically learn the visual features of the target without manually designing complex feature extraction algorithms. The confidence and category in the prediction information are used to determine the existence and type of the target, while the key point coordinates and flags are used for subsequent binding relationship prediction. The prediction of the bounding box ensures the precise positioning of the target and is the basis of the target detection task.

[0093] The target detection unit 30 is used to screen all the above prediction information corresponding to each of the above prediction targets to obtain the target prediction information corresponding to the above prediction target, and one above prediction target corresponds to one piece of the above target prediction information.

[0094] Specifically, after obtaining multiple prediction targets and their prediction information, it is necessary to filter this information and only retain the target prediction information with high confidence. The screening process usually includes setting a confidence threshold, retaining only the prediction information with confidence higher than the threshold, and performing non-maximum suppression on multiple prediction results of similar targets to eliminate overlapping prediction boxes. Each prediction target corresponds to one target prediction information. By setting the confidence threshold, low-quality detection results can be filtered out, improving the accuracy of the final binding relationship prediction.

[0095] The calculation unit 40 is used to perform bounding box matching calculation according to the target prediction information of each of the above-mentioned prediction targets to obtain a corresponding matching cost, where the above-mentioned matching cost is the Euclidean distance between the bounding box of the above-mentioned central prediction target and the bounding box of the above-mentioned potential binding prediction target.

[0096] Specifically, for the central prediction target and the potential binding prediction target, the Euclidean distance between the bounding box of the central prediction target and the bounding box of the potential binding prediction target is calculated as the matching cost. In the case of a one-to-one binding relationship, the matching cost is calculated as the average Euclidean distance between the center points of the two bounding boxes; in the case of a one-to-many binding relationship, the matching cost is the Euclidean distance from the center point of the central prediction target to the key points of other potential binding prediction targets. The Euclidean distance can quantify the spatial relationship between targets. The smaller the distance, the closer the two targets are, and the stronger the potential binding relationship. Through the matching cost, the algorithm can intelligently determine which potential binding prediction targets have a binding relationship with the central prediction target, thereby more accurately predicting the binding relationship and reducing the occurrence of misbinding and missed binding.

[0097] The matching unit 50 is used to perform matching using a matching algorithm according to the matching cost to obtain a binding relationship between the central prediction target and the potential binding prediction target.

[0098] Specifically, according to the calculated matching cost, a matching algorithm (such as a greedy algorithm, a depth-first search algorithm, or a Hungarian algorithm, etc.) is used to match and obtain the binding relationship between the central prediction target and the potential binding prediction target. The goal of the matching algorithm is to find a set of matching relationships so that the total cost of all matches is minimized or other optimization goals are met. The matching algorithm is a decision-making process for binding relationship prediction, which determines the optimal binding relationship based on the matching cost. Through the intelligent algorithm, the target pair with the minimum matching cost can be automatically found, thereby achieving efficient and accurate binding relationship prediction.

[0099] It should be noted that, in the face of overlapping targets, the intersection-and-union ratio between target frames may not necessarily reflect the binding relationship between targets (for example, the intersection-and-union ratio of the target frame of an electric vehicle and its rider may be smaller than the intersection-and-union ratio of the electric vehicle and the pedestrian next to it), resulting in a high misbinding rate in the traditional target post-binding algorithm, which further affects the effects of various intelligent algorithms, such as Figure 4 As shown, Figure 4A schematic diagram shows a mis-binding caused by overlapping targets bound by traditional methods. In the figure, A1 represents the key points of the head bounding box of predicted target A, A2 represents the key points of the human body bounding box of predicted target A, B1 represents the key points of the head bounding box of predicted target B, and B2 represents the key points of the human body bounding box of predicted target B. Mis-binding problems occur when binding by traditional methods. However, the method of the present invention can effectively identify binding relationships in which target frames do not overlap. For a binding relationship such as a pet dog and the owner who walks the dog, the two target frames may not overlap at all, causing the post-binding algorithm based on intersection-union ratio to be completely invalid, while the pre-target binding algorithm based on cross-target key points is basically unaffected. Figure 5 As shown, Figure 5 A schematic diagram of the effect of a cross-target key point pre-binding provided according to an embodiment of the present application is shown, in which A1 represents the key points of the head bounding box of the predicted target A, A2 represents the key points of the human body bounding box of the predicted target A, B1 represents the key points of the head bounding box of the predicted target B, and B2 represents the key points of the human body bounding box of the predicted target B. No mis-binding problem occurs when binding using the method of the present invention. The present invention can be applied to the recognition of various binding relationships such as face and human body, head and human body, vehicle and license plate, vehicle and driver, etc., effectively solving the mis-binding problem of the post-target binding algorithm in complex scenes, thereby improving the effect of the intelligent recognition algorithm in complex scenes.

[0100] In this embodiment, the first acquisition unit is used to acquire a target image, and the target image is an image obtained by preprocessing the original image; the extraction unit is used to extract features in the target image based on a deep neural network to obtain multiple prediction targets and multiple prediction information corresponding to the prediction targets. The prediction targets are divided into central prediction targets and potential binding prediction targets. The central prediction target is the subject to be predicted on the target image, and the potential binding prediction target is a target that has a potential interactive relationship with the central prediction target. Each of the prediction information includes the confidence, category, key point coordinates, key point flag and bounding box corresponding to the prediction target. The key point flag is used to predict whether the prediction target has key points, and the bounding box is used for box selection. The position of the predicted target in the target image, the key point represents the key position of the predicted target on the bounding box; the target detection unit is used to screen all the prediction information corresponding to each of the predicted targets, and obtain the target prediction information corresponding to the predicted target, one predicted target corresponds to one target prediction information; the calculation unit is used to match the bounding boxes according to the target prediction information of each of the predicted targets, and obtain the corresponding matching cost, and the matching cost is the Euclidean distance between the bounding box of the central predicted target and the bounding box of the potential binding prediction target; the matching unit is used to match according to the matching cost using the matching algorithm to obtain the binding relationship between the central predicted target and the potential binding prediction target. The key point of this application is cross-target, that is, it can predict the position of other targets (i.e., potential binding prediction targets) that have potential interaction relationships with the central prediction target, so that while identifying the categories and positions of the central prediction target and the potential binding prediction target, it is possible to predict the binding relationship between the targets. This method breaks through the existing limitation that target binding can only be performed through post-processing, so that the binding relationships such as subordination and interaction between targets can be intelligently predicted through deep learning models. Especially when the overlap of multiple targets is high, the binding relationship can be identified more stably and accurately, thereby improving the effect of each intelligent algorithm in complex scenarios. This application solves the problem in the prior art that the binding relationship between targets cannot be accurately identified when the overlap of multiple targets is high.

[0101] In order to improve the quality of image data, in an optional implementation manner, the first acquisition unit includes:

[0102] An acquisition module, used for acquiring the original image;

[0103] A format conversion module, used for converting the original image into a target image format to obtain a first image, wherein the target image format is at least an RGB format;

[0104] An adjustment module, used for adjusting the size of the first image to a target size and performing edge filling to obtain a second image;

[0105] The first normalization module is used to convert the second image into a tensor form and perform normalization processing to obtain the target image.

[0106] In the above embodiment, if Figure 3 As shown, the target front binding system of the present invention includes a pre-processing module, an image analysis module and a post-processing module. The function of the pre-processing module is to convert images with different resolutions, aspect ratios and compression protocols into a unified format required by the image analysis module. First, read the image file, which can be a static image directly read through the file system, or intercept a frame from the video stream. In addition, the original image may come from a variety of different devices, such as surveillance cameras, mobile phone cameras, etc., so their resolutions, aspect ratios and compression methods may be different. Convert the original image to RGB format or other formats to obtain the first image, because the RGB format is based on the color space of the red, green and blue channels, and the deep learning model usually trains RGB images, so it is necessary to convert the image data into a format that the model can understand. Adjust the image size to the input size required by the model by scaling, cropping or stretching, etc., usually this target size is a fixed size defined during model training. Then, fill the edges of the adjusted image to obtain the second image to avoid image information loss caused by image scaling. The image size is adjusted to meet the input requirements of the model, so that the model is based on the same size input when processing all images, thereby simplifying the calculation and improving the processing speed. Edge padding helps maintain the integrity of the image and avoids the loss of key information due to resizing, especially when the target is at the edge of the image. The second image is converted into a tensor, which is a common data structure in deep learning frameworks that can perform mathematical operations efficiently. Next, the pixel values ​​in the tensor are normalized, for example, the RGB pixel values ​​are scaled from the range of 0 to 255 to the range of 0 to 1, which helps with model training and prediction because the normalized data can speed up model convergence and reduce numerical instability. It ensures that the image data has been standardized and optimized before entering the deep learning model, so that the model can perform target detection and binding relationship prediction more effectively. This not only improves the computational efficiency of the model, but also enhances the generalization ability of the model under different image conditions, which is especially important for processing image data in complex scenes.

[0107] In order to improve the accuracy and stability of target detection and analysis, in an optional embodiment, the device further includes:

[0108] A second acquisition unit is used to obtain an initial sample target detection data set before extracting features in the target image based on a deep neural network to obtain multiple predicted targets and multiple prediction information corresponding to the predicted targets, and to enclose and annotate sample binding relationships for each sample in the initial sample target detection data set, wherein the initial sample target detection data set includes multiple sample targets, and each of the sample targets includes a sample category, a sample bounding box, and a sample confidence;

[0109] A generating unit, configured to generate sample key points and sample key point flags of the corresponding sample bounding boxes according to the sample binding relationships, and to update the initial sample target detection data set according to the sample key points and the sample key point flags to obtain a sample target detection data set;

[0110] A training unit, used to execute a training step, input the sample targets in the sample target detection data set into an initial deep neural network, and iteratively train the initial deep neural network using a target loss function to obtain a current deep neural network, wherein the initial deep neural network is an untrained deep neural network, and the target loss function at least includes a key point landmark classification loss function and a key point coordinate regression loss function;

[0111] A testing unit, used to execute a testing step, test the current deep neural network using a test set, and obtain a testing index, wherein the testing index at least includes an accuracy rate and a false alarm rate;

[0112] The repeating unit is used to execute the repetitive steps. When the above test indicators do not meet the set accuracy, the sample target is increased or the hyperparameter is adjusted, and the above training steps and the above test steps are repeatedly executed in sequence at least once until the above test indicators meet the set accuracy.

[0113] In the above embodiments, usually, the target detection data set does not have the annotation of the binding relationship, and the method proposed in this patent requires the model to predict the target category, bounding box and binding relationship at the same time, so the target detection data set (including training and test sets) is required. First, an initial target detection data set is obtained from an existing database or acquisition system. This data set contains multiple image samples, and each sample has the target category, bounding box and confidence information marked. Then, for the targets in each sample, the binding relationship between them is manually or using auxiliary software. For example, in an image containing multiple human bodies and human heads, it is necessary to mark the binding relationship between each head frame and the corresponding human body frame. Basic data for learning binding relationships is provided, so that the model can learn the association between target detection and binding relationship prediction, so that accurate binding prediction can be performed on unknown data. Based on the marked binding relationship, a key point is generated for each target frame, and this key point points to the center of other targets bound to it. At the same time, a key point flag is generated for each target frame to indicate whether the target frame needs to be bound. Then, these key point coordinates and flag information are added to the sample data set to update the original data set. The generation of keypoints and keypoint landmarks is a key data preparation process for the pre-training binding algorithm model. Keypoints provide a direct indication of the binding relationship between objects, while landmarks help the model understand which object boxes need to be bound. The updated sample object detection dataset is input into the initial deep neural network, which can be YOLO, Faster R-CNN, or any other deep learning architecture suitable for object detection. Then, the target loss function is defined, which includes at least the keypoint landmark classification loss function (such as BCELoss, BinaryCrossEntropyLoss, Binary Cross Entropy Loss) and the keypoint coordinate regression loss function (such as L2Loss, MeanSquared Error Loss, Mean Squared Error Loss). In each iteration, the model predicts the object bounding box, category, keypoint coordinates, and keypoint landmarks compared with the true value of the sample, calculates the loss, and then uses the optimizer (such as Adam) to update the network parameters to minimize the loss function. Through iterative training, the model gradually learns to predict the category and location of the object from the input image, and can also predict the keypoint coordinates and landmarks, and then predict the binding relationship between two object boxes. The definition of the loss function ensures that the model can optimize the performance of both object detection and binding relationship prediction, thereby improving the accuracy and robustness of the overall model. After the model is trained, it is evaluated using an independent test set that was not involved in the training. The test set contains image samples that are also labeled with object categories, bounding boxes, and binding relationships.The model predicts the images in the test set, outputs the predicted target box, category, key point coordinates and key point flags, and then compares them with the true values ​​to calculate the predicted accuracy and false alarm rate and other indicators. Indicators such as accuracy and false alarm rate can reflect the performance of the model in target detection and binding relationship prediction, and help determine whether the model has reached the expected accuracy standard. If the test indicators do not meet the preset accuracy standards, it is necessary to further optimize the model by adding more sample data or adjusting the model's hyperparameters (such as learning rate, regularization term). Specifically, samples containing complex binding relationships can be added to improve the robustness of the model, or the learning rate can be adjusted, the number of training rounds can be increased, etc., to improve the training effect of the model. Then, the above training steps and the above test steps are repeated in sequence until the test indicators of the model meet the set accuracy, and finally a deep neural network model that performs well in both target detection and binding relationship prediction is obtained. Through this process, the model can learn and predict complex binding relationships, thereby realizing a more accurate and efficient front binding algorithm in target detection tasks, especially in complex scenes with high target overlap, which can significantly improve the accuracy and stability of target detection and analysis.

[0114] In order to improve the accuracy of the detection result, in an optional implementation manner, the target detection unit includes:

[0115] A mapping processing module, used to map the confidence in each of the above prediction information through a Sigmoid activation function to obtain a corresponding target confidence, where the target confidence is a value between 0 and 1;

[0116] A second normalization module, used to normalize the categories in each of the above prediction information through a normalized exponential function to obtain a corresponding target category vector and a target category, wherein the elements in the above target category vector are all values ​​between 0 and 1 and the sum of all the above elements is equal to 1;

[0117] A first retention module is used to retain the potential binding prediction targets whose target confidence is greater than or equal to the confidence threshold, to obtain binding prediction targets;

[0118] A coordinate conversion module, used for converting the coordinates of the bounding box in all the prediction information corresponding to the bounding prediction target and the central prediction target from local coordinates to global coordinates on the target image;

[0119] The second retention module is used to adopt a non-maximum suppression algorithm to retain the prediction information with the highest confidence for each of the above-mentioned prediction targets, and obtain the target prediction information corresponding to each of the above-mentioned prediction targets.

[0120] In the above embodiment, if Figure 4As shown in the figure, the main function of the post-processing module is to convert the output of the image analysis module into the required prediction values, such as confidence, category, bounding box, binding relationship, etc. The post-processing module is divided into a target detection post-processing module and a binding relationship post-processing module. The output of the deep neural network usually contains the original confidence value (usually ranging from negative infinity to positive infinity). These values ​​need to be converted through the Sigmoid activation function to map them between 0 and 1, so that the target confidence obtained can represent the model's probability estimate of the existence of a certain target in the image. The categories in each of the above prediction information are normalized by the normalized exponential function to obtain the corresponding target category vector and target category implementation process: The category information output by the model is usually a set of original category scores. By using the normalized exponential function (usually referring to the Softmax function), these scores are converted into probability vectors, where the sum of the probabilities of all categories is equal to 1. Based on this probability vector, the category with the highest probability can be selected as the final target category. Based on the obtained target confidence, a confidence threshold is set to retain only those potential binding targets with confidence higher than the threshold. Function and effect: Confidence threshold filtering is a key step to remove low-quality detection frames. By setting a reasonable threshold, detection results that the model is not sure about can be excluded, improving the accuracy and reliability of the detected target. The predictions of deep learning models are usually given in the form of local coordinates, that is, the coordinates relative to the grid cell to which the detection frame belongs. In order to accurately mark the position of the detection frame on the original image, it is necessary to convert these local coordinates into global coordinates, that is, the actual coordinate position on the image. Coordinate conversion ensures the correct position of the detection frame on the original image, which is crucial for the visualization of target detection and binding relationships and subsequent processing. The non-maximum suppression (NMS) algorithm is used to filter out those overlapping bounding boxes. Specifically, for each predicted target, the non-maximum suppression algorithm compares the confidence of all bounding boxes of the predicted target, retains the box with the highest confidence and low overlap with other boxes, and removes the rest. The non-maximum suppression algorithm solves the problem of multiple detections of the same target, ensuring that each target is detected only once and the detection result has the highest confidence. It realizes the effective transformation from model output to target detection and binding relationship prediction results. The conversion of confidence and category probability, confidence threshold filtering, coordinate transformation, and the application of the NMS algorithm together constitute a complete post-processing process, which aims to optimize the accuracy of the detection results, reduce redundancy, and improve the robustness of the model. This ensures that the model can not only accurately detect the target, but also effectively identify the binding relationship between the targets, especially in the case of dense targets and high overlap in the image, which can significantly improve the efficiency and accuracy of the overall detection and analysis.

[0121] In order to significantly reduce false positives and false negatives in binding relationship prediction, in an optional implementation manner, the device further includes:

[0122] A determination unit is used to screen all the prediction information corresponding to each of the prediction targets to obtain the target prediction information corresponding to the prediction targets. If there is only one binding prediction target and the key point flag of the central prediction target is 0, determine that the central prediction target has no binding relationship in the target image. The binding prediction target is the potential binding prediction target whose target confidence is greater than or equal to a confidence threshold. The target confidence is the confidence after mapping the confidence in the prediction information.

[0123] In the above embodiment, multiple potential binding prediction targets have been obtained, which are prediction targets whose target confidence is greater than or equal to the confidence threshold. These targets are candidate targets that the model believes exist in the current image and may have a binding relationship with other targets. In a one-to-one binding task, as long as there is a binding relationship, the key point flag must be 1. If the key point flag is 0, it means that the model has not predicted the binding key points related to the central target, indicating that the target may be independent in the current scene and does not require a binding relationship. Therefore, it is confirmed that in the current scene, the model only predicts one type of binding target. For example, in the binding scene of the human body and the human head, if the model only predicts the human head as a potential binding target, and the key point flag of the central prediction target is 0, then it can be determined that the central prediction target has no binding relationship in the target image. The false positives and false negatives of the binding relationship prediction are significantly reduced, and the overall accuracy of the binding relationship recognition is improved.

[0124] It should be noted that for a one-to-one target binding task, the key point is defined as follows: if bounding box A and bounding box B are a pair of binding relationships, then the key point of A is the center point of B, and the key point of B is the center point of A; for a one-to-many target binding task, the key point is defined as follows: if bounding box A and bounding box B, bounding box A and bounding box C both form a binding relationship, then A has no key point, and the key points of B and C are the center point of A.

[0125] In order to accurately determine the binding relationship between two targets, in an optional implementation, the calculation unit includes:

[0126] A first calculation module is used to calculate the Euclidean distance between the key points of the bounding box of each of the bounding prediction targets and the center point of the bounding box of the central prediction target in the presence of multiple bounding prediction targets, to obtain the corresponding matching cost, wherein the bounding prediction target is the potential bounding prediction target whose target confidence is greater than or equal to a confidence threshold, the target confidence is the confidence after mapping the confidence in the prediction information, and the center point is the center position of the bounding box;

[0127] The second calculation module is used to calculate the average of the first Euclidean distance and the second Euclidean distance to obtain the matching cost when there is only one of the above-mentioned binding prediction targets and the key point flag of the above-mentioned central prediction target is 1, wherein the first Euclidean distance is the Euclidean distance between the above-mentioned key point of the above-mentioned enclosing box of the above-mentioned binding prediction target and the above-mentioned center point of the above-mentioned enclosing box of the above-mentioned central prediction target, and the second Euclidean distance is the Euclidean distance between the above-mentioned key point of the above-mentioned enclosing box of the above-mentioned central prediction target and the above-mentioned center point of the above-mentioned enclosing box of the above-mentioned binding prediction target.

[0128] In the above embodiment, all high-confidence potential binding prediction targets (confidence greater than or equal to the threshold) are obtained. These prediction targets are candidate targets that the model believes exist in the current image and may be bound to the central target. There are multiple binding prediction targets, which means it is a one-to-many binding task. For example, the binding between the three bounding boxes of the face, head, and body of the same person. For the central prediction target and each binding prediction target, the Euclidean distance between the key point of the binding prediction target and the center point of the central prediction target is calculated. The key point here is the point predicted by the deep learning network model in the detection phase to indicate the potential binding relationship, and the center point is the geometric center of the target bounding box. The calculated Euclidean distance is used as the matching cost to quantify the strength of the binding relationship between the central prediction target and the binding prediction target. The smaller the Euclidean distance, the stronger the binding relationship between the binding prediction target and the central prediction target, and the lower the matching cost. When there is only one binding prediction target (for example, the head frame is used as the binding target of the human frame) and the key point flag of the central prediction target (such as the human frame) is 1, that is, the model predicts that the central target needs to be bound. First, the first Euclidean distance from the key point of the binding prediction target to the center point of the central prediction target is calculated. Then, the second Euclidean distance from the key point of the central prediction target to the center point of the binding prediction target is calculated. The average of the first Euclidean distance and the second Euclidean distance is used as the final matching cost. This step takes into account the bidirectionality of the binding relationship and ensures that the judgment of the binding relationship is more accurate and fair. By calculating the distance between the key point and the center point as the matching cost, the binding relationship between the two targets can be judged more accurately, avoiding the limitations of the traditional method based on intersection and union ratio in complex scenes with overlapping targets. The calculation method of the Euclidean distance is not affected by the size and shape of the target frame, making the algorithm more stable when processing targets of different sizes and shapes, and reducing the possibility of misjudgment. This key point-based matching cost calculation strategy not only enhances the performance of the algorithm in complex scenes, but also simplifies the matching process and improves the processing efficiency of the algorithm.

[0129] In order to significantly enhance the reliability of binding relationship prediction, in an optional implementation manner, the matching unit includes:

[0130] A first matching module, configured to perform matching by using a first matching algorithm to obtain a first binding relationship, wherein the first matching algorithm is at least one of the following: a greedy algorithm, a depth-first search algorithm, and a Hungarian algorithm;

[0131] A second matching module, configured to perform matching by using a second matching algorithm to obtain a second binding relationship, wherein the second matching algorithm is the matching algorithm other than the first matching algorithm;

[0132] A first determining module, configured to determine the first binding relationship as a binding relationship between the central prediction target and the potential binding prediction target when the first binding relationship is the same as the second binding relationship;

[0133] A third matching module is used to use a third matching algorithm to perform matching to obtain a third binding relationship when the first binding relationship is different from the second binding relationship, and the third matching algorithm is the matching algorithm other than the first matching algorithm and the second matching algorithm;

[0134] The second determination module is used to determine the third binding relationship as the binding relationship between the central prediction target and the potential binding prediction target when the third binding relationship is the same as the first binding relationship or the third binding relationship is the same as the second binding relationship.

[0135] In the above embodiment, a first matching algorithm is used, usually one of a greedy algorithm, a depth-first search algorithm or a Hungarian algorithm, to make a preliminary judgment on the match between the central prediction target and the potential binding prediction target. These algorithms determine the best binding relationship based on the matching cost (such as Euclidean distance). The greedy algorithm may select the target pair with the smallest matching cost as the binding relationship; the depth-first search algorithm may explore all possible binding combinations and select the best one; the Hungarian algorithm is designed to solve the maximum weight matching in the bipartite graph, and it can find a set of matching relationships so that the total cost of all matches (i.e., the total distance) is minimized. The second matching algorithm is used for matching to obtain a second possible binding relationship. The second matching algorithm should be a choice other than the first matching algorithm. For example, if the first algorithm is a greedy algorithm, the second algorithm can be a depth-first search algorithm or a Hungarian algorithm. Compare the first binding relationship and the second binding relationship. If the binding relationships obtained by the two algorithms are the same, that is, they both point to the same set of target pairs, then it can be determined that this is the correct binding relationship between the central prediction target and the potential binding prediction target. If the binding relationships obtained by the first algorithm and the second algorithm are different, it means that there are different results caused by uncertainty or algorithm differences. At this time, a third matching algorithm is needed to perform further matching to obtain a third possible binding relationship. Compare the third binding relationship with the first binding relationship or the second binding relationship. If the result of the third algorithm is the same as the result of any previous algorithm, then this result is adopted as the binding relationship between the central prediction target and the potential binding prediction target. If the first binding relationship, the second binding relationship, and the third binding relationship are not the same, a prompt is issued that there may be no binding relationship, or a detection error problem has occurred. By adopting multiple matching algorithms and comparing the consistency of the results, the reliability of the binding relationship prediction can be significantly enhanced. If different algorithms produce consistent results, this indicates that the prediction results are robust and reduce the errors that may be caused by a single algorithm. Using different matching algorithms can cope with different types of binding relationships and scene complexities. For example, the greedy algorithm is suitable for simple situations, while the Hungarian algorithm is more suitable for complex matching problems. This diversity ensures that the algorithm can flexibly cope with various scenarios and improves the overall robustness. When the results of the initial two algorithms are inconsistent, the third algorithm is introduced to help resolve the conflict, reduce the uncertainty caused by algorithm selection, and improve the certainty of the binding relationship prediction. Although the use of multiple algorithms may increase the amount of calculation, in most cases, the algorithm in the first step can give reliable results and reduce the dependence on more complex algorithms. This strategy optimizes the use of computing resources as much as possible while ensuring accuracy.

[0136] The target pre-binding device based on cross-target key points includes a processor and a memory. The first acquisition unit, extraction unit, target detection unit, etc. are all stored in the memory as program units, and the processor executes the program units stored in the memory to implement corresponding functions. The modules are all located in the same processor; or, the modules are located in different processors in any combination.

[0137] The processor includes a kernel, which retrieves the corresponding program unit from the memory. One or more kernels can be set, and the problem of the inability to accurately identify the binding relationship between targets when the overlap of multiple targets is high in the prior art can be solved by adjusting the kernel parameters.

[0138] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.

[0139] An embodiment of the present invention provides a computer-readable storage medium, which includes a stored program, wherein when the program is running, the device where the computer-readable storage medium is located is controlled to execute the target pre-binding method based on cross-target key points.

[0140] Specifically, the target front binding method based on cross-target key points includes:

[0141] Step S201, obtaining a target image, wherein the target image is an image obtained by preprocessing the original image;

[0142] Step S202, extracting features in the target image based on a deep neural network, obtaining multiple prediction targets and multiple prediction information corresponding to the prediction targets, wherein the prediction targets are divided into a central prediction target and a potential binding prediction target, wherein the central prediction target is a subject to be predicted on the target image, and the potential binding prediction target is a target having a potential interactive relationship with the central prediction target, and each of the prediction information includes a confidence level, a category, key point coordinates, a key point flag, and a bounding box corresponding to the prediction target, wherein the key point flag is used to predict whether the prediction target has a key point, and the bounding box is used to select the position of the prediction target in the target image, and the key point represents a key position of the prediction target on the bounding box;

[0143] Step S203, screening all the prediction information corresponding to each of the prediction targets to obtain target prediction information corresponding to the prediction target, where one prediction target corresponds to one target prediction information;

[0144] Step S204, performing bounding box matching calculation according to the target prediction information of each of the predicted targets to obtain a corresponding matching cost, where the matching cost is the Euclidean distance between the bounding box of the central predicted target and the bounding box of the potential bounding predicted target;

[0145] Step S205 , performing matching using a matching algorithm according to the matching cost, and obtaining a binding relationship between the central prediction target and the potential binding prediction target.

[0146] An embodiment of the present invention provides a processor, which is used to run a program, wherein the program executes the target pre-binding method based on cross-target key points when running.

[0147] An embodiment of the present invention provides a target pre-binding system, the target pre-binding system includes a processor, a memory, and a program stored in the memory and executable on the processor, and when the processor executes the program, at least the following steps are implemented:

[0148] Step S201, obtaining a target image, wherein the target image is an image obtained by preprocessing the original image;

[0149] Step S202, extracting features in the target image based on a deep neural network, obtaining multiple prediction targets and multiple prediction information corresponding to the prediction targets, wherein the prediction targets are divided into a central prediction target and a potential binding prediction target, wherein the central prediction target is a subject to be predicted on the target image, and the potential binding prediction target is a target having a potential interactive relationship with the central prediction target, and each of the prediction information includes a confidence level, a category, key point coordinates, a key point flag, and a bounding box corresponding to the prediction target, wherein the key point flag is used to predict whether the prediction target has a key point, and the bounding box is used to select the position of the prediction target in the target image, and the key point represents a key position of the prediction target on the bounding box;

[0150] Step S203, screening all the prediction information corresponding to each of the prediction targets to obtain target prediction information corresponding to the prediction target, where one prediction target corresponds to one target prediction information;

[0151] Step S204, performing bounding box matching calculation according to the target prediction information of each of the predicted targets to obtain a corresponding matching cost, where the matching cost is the Euclidean distance between the bounding box of the central predicted target and the bounding box of the potential bounding predicted target;

[0152] Step S205 , performing matching using a matching algorithm according to the matching cost, and obtaining a binding relationship between the central prediction target and the potential binding prediction target.

[0153] The present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing a program for initializing at least the following method steps:

[0154] Step S201, obtaining a target image, wherein the target image is an image obtained by preprocessing the original image;

[0155] Step S202, extracting features in the target image based on a deep neural network, obtaining multiple prediction targets and multiple prediction information corresponding to the prediction targets, wherein the prediction targets are divided into a central prediction target and a potential binding prediction target, wherein the central prediction target is a subject to be predicted on the target image, and the potential binding prediction target is a target having a potential interactive relationship with the central prediction target, and each of the prediction information includes a confidence level, a category, key point coordinates, a key point flag, and a bounding box corresponding to the prediction target, wherein the key point flag is used to predict whether the prediction target has a key point, and the bounding box is used to select the position of the prediction target in the target image, and the key point represents a key position of the prediction target on the bounding box;

[0156] Step S203, screening all the prediction information corresponding to each of the prediction targets to obtain target prediction information corresponding to the prediction target, where one prediction target corresponds to one target prediction information;

[0157] Step S204, performing bounding box matching calculation according to the target prediction information of each of the predicted targets to obtain a corresponding matching cost, where the matching cost is the Euclidean distance between the bounding box of the central predicted target and the bounding box of the potential bounding predicted target;

[0158] Step S205 , performing matching using a matching algorithm according to the matching cost, and obtaining a binding relationship between the central prediction target and the potential binding prediction target.

[0159] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order than here, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.

[0160] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0161] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0162] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0163] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0164] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0165] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0166] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0167] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0168] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:

[0169] 1) The target pre-binding method based on cross-target key points of the present application, first, obtain a target image, the target image is an image after preprocessing the original image; then, based on the deep neural network, extract the features in the target image, obtain multiple prediction targets and multiple prediction information corresponding to the prediction targets, the prediction targets are divided into central prediction targets and potential binding prediction targets, the central prediction target is the subject to be predicted on the target image, the potential binding prediction target is a target that has a potential interactive relationship with the central prediction target, each of the prediction information includes the confidence, category, key point coordinates, key point flag and bounding box corresponding to the prediction target, the key point flag is used to predict whether the prediction target has a key point, The above-mentioned bounding box is used to select the position of the above-mentioned predicted target in the above-mentioned target image, and the above-mentioned key point represents the key position of the above-mentioned predicted target on the above-mentioned bounding box; then, all the above-mentioned prediction information corresponding to each of the above-mentioned predicted targets is screened to obtain the target prediction information corresponding to the above-mentioned predicted target, and one above-mentioned predicted target corresponds to one above-mentioned target prediction information; then, the bounding box matching calculation is performed according to the above-mentioned target prediction information of each of the above-mentioned predicted targets to obtain the corresponding matching cost, and the above-mentioned matching cost is the Euclidean distance between the bounding box of the above-mentioned central predicted target and the bounding box of the above-mentioned potential binding prediction target; finally, the matching is performed using the matching algorithm according to the above-mentioned matching cost to obtain the binding relationship between the above-mentioned central predicted target and the above-mentioned potential binding prediction target. The key point of this application is cross-target, that is, it can predict the position of other targets (i.e., potential binding prediction targets) that have potential interactive relationships with the central prediction target, so that while identifying the categories and positions of the central prediction target and the potential binding prediction target, the binding relationship between the targets can be predicted. This method breaks through the existing limitation that target binding can only be performed through post-processing, so that the binding relationships such as subordination and interaction between targets can be intelligently predicted through deep learning models. Especially when the overlap of multiple targets is high, the binding relationship can be identified more stably and accurately, thereby improving the effect of each intelligent algorithm in complex scenarios. This application solves the problem in the prior art that the binding relationship between targets cannot be accurately identified when the overlap of multiple targets is high.

[0170] 2) The target pre-binding device based on cross-target key points of the present application comprises a first acquisition unit for acquiring a target image, wherein the target image is an image obtained by preprocessing the original image; an extraction unit for extracting features in the target image based on a deep neural network to obtain multiple prediction targets and multiple prediction information corresponding to the prediction targets. The prediction targets are divided into a central prediction target and a potential binding prediction target. The central prediction target is a subject to be predicted on the target image, and the potential binding prediction target is a target having a potential interactive relationship with the central prediction target. Each of the prediction information comprises a confidence level, a category, key point coordinates, a key point flag, and a bounding box corresponding to the prediction target. The key point flag is used to predict whether the prediction target has a key point. The enclosing box is used to select the position of the predicted target in the target image, and the key point represents the key position of the predicted target on the enclosing box; the target detection unit is used to screen all the prediction information corresponding to each of the predicted targets to obtain the target prediction information corresponding to the predicted target, and one predicted target corresponds to one target prediction information; the calculation unit is used to match the enclosing boxes according to the target prediction information of each of the predicted targets to obtain the corresponding matching cost, and the matching cost is the Euclidean distance between the enclosing box of the central predicted target and the enclosing box of the potential binding prediction target; the matching unit is used to match according to the matching cost using the matching algorithm to obtain the binding relationship between the central predicted target and the potential binding prediction target. The key point of this application is cross-target, that is, it can predict the position of other targets (i.e., potential binding prediction targets) that have potential interactive relationships with the central prediction target, so that while identifying the categories and positions of the central prediction target and the potential binding prediction target, it is possible to predict the binding relationship between the targets. This method breaks through the existing limitation that target binding can only be performed through post-processing, so that the binding relationships such as subordination and interaction between targets can be intelligently predicted through deep learning models. Especially when the overlap of multiple targets is high, the binding relationship can be identified more stably and accurately, thereby improving the effect of each intelligent algorithm in complex scenarios. This application solves the problem in the prior art that the binding relationship between targets cannot be accurately identified when the overlap of multiple targets is high.

[0171] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A target front binding method based on cross-target key points, characterized in that: include: Acquire a target image, wherein the target image is an image obtained by preprocessing the original image; Extract features in the target image based on a deep neural network, obtain multiple prediction targets and multiple prediction information corresponding to the prediction targets, the prediction targets are divided into a central prediction target and a potential binding prediction target, the central prediction target is the subject to be predicted on the target image, the potential binding prediction target is a target that has a potential interactive relationship with the central prediction target, each of the prediction information includes the confidence, category, key point coordinates, key point flag and bounding box corresponding to the prediction target, the key point flag is used to predict whether the prediction target has a key point, the bounding box is used to frame the position of the prediction target in the target image, and the key point represents the key position of the prediction target on the bounding box; Screening all the prediction information corresponding to each prediction target to obtain target prediction information corresponding to the prediction target, where one prediction target corresponds to one target prediction information; Performing bounding box matching calculation according to the target prediction information of each of the predicted targets to obtain a corresponding matching cost, wherein the matching cost is the Euclidean distance between the bounding box of the central predicted target and the bounding box of the potential bounding predicted target; Matching is performed using a matching algorithm according to the matching cost to obtain a binding relationship between the central prediction target and the potential binding prediction target.

2. The method according to claim 1, characterized in that Get the target image, including: Get the original image; Converting the original image into a target image format to obtain a first image, wherein the target image format is at least an RGB format; Adjusting the size of the first image to a target size and performing edge filling to obtain a second image; The second image is converted into a tensor form and normalized to obtain the target image.

3. The method according to claim 1, characterized in that Before extracting features from the target image based on a deep neural network to obtain a plurality of prediction targets and a plurality of prediction information corresponding to the prediction targets, the method further includes: Acquire an initial sample target detection data set, and annotate a sample binding relationship for each sample in the initial sample target detection data set, wherein the initial sample target detection data set includes a plurality of sample targets, and each of the sample targets includes a sample category, a sample bounding box, and a sample confidence; Generate the sample key points and sample key point flags of the corresponding sample bounding boxes according to each of the sample binding relationships, and update the initial sample target detection data set according to the sample key points and the sample key point flags to obtain a sample target detection data set; A training step, inputting the sample targets in the sample target detection data set into an initial deep neural network, and iteratively training the initial deep neural network using a target loss function to obtain a current deep neural network, wherein the initial deep neural network is an untrained deep neural network, and the target loss function at least includes a key point landmark classification loss function and a key point coordinate regression loss function; A testing step, using a test set to test the current deep neural network to obtain a test index, wherein the test index at least includes an accuracy rate and a false alarm rate; Repeating the steps, when the test index does not meet the set accuracy, increasing the sample target or adjusting the hyperparameters, and repeating the training step and the testing step in sequence at least once until the test index meets the set accuracy.

4. The method according to claim 1, characterized in that: Screening all the prediction information corresponding to each prediction target to obtain target prediction information corresponding to the prediction target includes: The confidence in each prediction information is mapped by a Sigmoid activation function to obtain a corresponding target confidence, where the target confidence is a value between 0 and 1; Normalizing the categories in each of the prediction information by a normalized exponential function to obtain a corresponding target category vector and a target category, wherein the elements in the target category vector are all values ​​between 0 and 1 and the sum of all the elements is equal to 1; Retain the potential binding prediction targets whose target confidence is greater than or equal to the confidence threshold to obtain binding prediction targets; Convert the coordinates of the bounding boxes in all the prediction information corresponding to the bounding prediction target and the central prediction target from local coordinates to global coordinates on the target image; A non-maximum suppression algorithm is used to retain the prediction information with the highest target confidence for each prediction target, so as to obtain the target prediction information corresponding to each prediction target.

5. The method according to claim 1, characterized in that After screening all the prediction information corresponding to each prediction target to obtain the target prediction information corresponding to the prediction target, the method further includes: When there is only one binding prediction target and the key point flag of the central prediction target is 0, it is determined that the central prediction target has no binding relationship in the target image. The binding prediction target is the potential binding prediction target whose target confidence is greater than or equal to the confidence threshold, and the target confidence is the confidence after mapping the confidence in the prediction information.

6. The method according to claim 1, characterized in that Performing bounding box matching calculations based on the target prediction information of each predicted target to obtain corresponding matching costs includes: In the case where there are multiple binding prediction targets, the Euclidean distance between the key points of the bounding box of each binding prediction target and the center point of the bounding box of the central prediction target is calculated to obtain the corresponding matching cost, the binding prediction target is the potential binding prediction target whose target confidence is greater than or equal to the confidence threshold, the target confidence is the confidence after mapping the confidence in the prediction information, and the center point is the center position of the bounding box; When there is only one binding prediction target and the key point flag of the central prediction target is 1, the average of the first Euclidean distance and the second Euclidean distance is calculated to obtain the matching cost, wherein the first Euclidean distance is the Euclidean distance between the key point of the bounding box of the binding prediction target and the center point of the bounding box of the central prediction target, and the second Euclidean distance is the Euclidean distance between the key point of the bounding box of the central prediction target and the center point of the bounding box of the binding prediction target.

7. The method according to claim 1, characterized in that The matching algorithm includes at least a greedy algorithm, a depth-first search algorithm and a Hungarian algorithm. The matching algorithm is used to perform matching according to the matching cost to obtain a binding relationship between the central prediction target and the potential binding prediction target, including: Using a first matching algorithm to perform matching to obtain a first binding relationship, wherein the first matching algorithm is at least one of the following: a greedy algorithm, a depth-first search algorithm, and a Hungarian algorithm; A second matching algorithm is used for matching to obtain a second binding relationship, where the second matching algorithm is the matching algorithm other than the first matching algorithm; In a case where the first binding relationship is the same as the second binding relationship, determining the first binding relationship as a binding relationship between the central prediction target and the potential binding prediction target; In the case that the first binding relationship is different from the second binding relationship, a third matching algorithm is used for matching to obtain a third binding relationship, and the third matching algorithm is the matching algorithm other than the first matching algorithm and the second matching algorithm; In the case where the third binding relationship is the same as the first binding relationship or the third binding relationship is the same as the second binding relationship, the third binding relationship is determined as the binding relationship between the central prediction target and the potential binding prediction target.

8. A target pre-binding device based on cross-target key points, characterized in that: The device comprises: A first acquisition unit, used to acquire a target image, wherein the target image is an image obtained by preprocessing the original image; An extraction unit is used to extract features in the target image based on a deep neural network to obtain multiple prediction targets and multiple prediction information corresponding to the prediction targets, wherein the prediction targets are divided into a central prediction target and a potential binding prediction target, wherein the central prediction target is a subject to be predicted on the target image, and the potential binding prediction target is a target that has a potential interactive relationship with the central prediction target, and each prediction information includes a confidence, category, key point coordinates, key point flag, and bounding box corresponding to the prediction target, wherein the key point flag is used to predict whether the prediction target has a key point, and the bounding box is used to select the position of the prediction target in the target image, and the key point represents a key position of the prediction target on the bounding box; A target detection unit, used for screening all the prediction information corresponding to each prediction target to obtain target prediction information corresponding to the prediction target, wherein one prediction target corresponds to one target prediction information; A calculation unit, configured to perform bounding box matching calculation according to the target prediction information of each predicted target to obtain a corresponding matching cost, wherein the matching cost is a Euclidean distance between the bounding box of the central predicted target and the bounding box of the potential bounding predicted target; A matching unit is used to perform matching using a matching algorithm according to the matching cost to obtain a binding relationship between the central prediction target and the potential binding prediction target.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 7.

10. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.