Target matching method, system and device for multi-target scene and medium

By constructing a cost matrix that integrates spatial distance and prior constraints of human structure for global optimal matching, the problem of multiple matching of one object and mismatch of proximity in human-object interaction behavior recognition in complex multi-person scenarios is solved, thereby improving the accuracy and reliability of recognition.

CN121811299APending Publication Date: 2026-04-07TIANSHU TONGYANG (ZHEJIANG) SEMICONDUCTOR TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-02
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing human-object interaction behavior recognition technologies struggle to achieve accuracy and reliability in complex scenarios involving multiple people. This is primarily because existing methods treat the matching problem as a local rule judgment, leading to frequent instances of multiple matching of the same object or mismatching of nearby objects, failing to balance accuracy, stability, and scalability.

Method used

A cost matrix is ​​constructed, which integrates spatial distance constraints and prior constraints of human body structure. The solution is obtained through global optimal matching, and unreasonable matching is filtered by a preset cost threshold, thereby improving the accuracy and robustness of the matching results.

Benefits of technology

It effectively avoids mismatches caused by local rule matching, and improves the accuracy and reliability of target association in multi-target scenarios, especially the recognition effect under dense occlusion and complex poses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121811299A_ABST
    Figure CN121811299A_ABST
Patent Text Reader

Abstract

The invention provides a target matching method, system and device for a multi-target scene and a medium, and the method comprises the steps: obtaining detection information of a plurality of human body targets and a plurality of to-be-associated targets in a monitoring image; constructing a cost matrix based on the detection information; solving the cost matrix to obtain an optimal matching result of the human body target and the to-be-associated target; based on the optimal matching result and a preset cost threshold value, unreasonable matching is filtered, and a final matching result is obtained. According to the method, a target association problem in a multi-person and multi-target scene is reconstructed into a global optimal matching problem, the global optimal solution is obtained by constructing the cost matrix and solving, and the phenomenon of one-target multi-matching or nearby mismatching caused by local rule matching is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of target matching technology, and more specifically, to a target matching method, system, device, and medium for multi-target scenarios. Background Technology

[0002] With the continuous development of video surveillance systems, intelligent security systems, and smart city construction, video image-based human behavior recognition technology has been widely applied in public safety management, industrial safety supervision, park management, and transportation. Among these, human-object interaction behavior recognition, as a typical task for identifying uncivilized or safety hazard behaviors, has significant application value in specific monitored locations (such as factory workshops, gas stations, warehouses, hospitals, subway stations, and schools). In practical applications, human-object interaction behavior recognition typically relies on video images acquired by front-end camera equipment. The back-end intelligent analysis system processes the images or video streams. When multiple people appear simultaneously in the monitored scene, or even when multiple people interact with the target to be associated simultaneously, the system needs to accurately associate the target with the corresponding human body target to determine which specific person exhibited the specific interactive behavior, thereby achieving accurate alarms, locating the responsible party, and subsequent management.

[0003] However, current mainstream technologies for human-object interaction recognition generally employ matching methods based on simple spatial rules. For example, after detecting a human target and a target to be associated, the center point of the target to be associated is linked to the nearest human target, or association is made using a fixed distance threshold or area overlap ratio (IoU). Essentially, these methods treat the matching problem between the human body and the target to be associated as a local rule-based problem, rather than a global optimization matching problem. Without unified mathematical modeling, matching strategies rely on empirical thresholds or heuristic rules, making it difficult to simultaneously ensure accuracy, stability, and scalability in complex scenarios. When multiple human targets and multiple targets to be associated exist in the monitoring footage, matching methods based on nearest distance or fixed thresholds are prone to mismatching a target to be associated with someone other than the actual person interacting with it, leading to multiple matchings of the same object or mismatches based on proximity. This results in incorrect location of the behavioral subject and reduces the accuracy and reliability of human-object interaction recognition in complex scenarios with multiple people. Summary of the Invention

[0004] The purpose of this application is to provide a target matching method, system, device, and medium for multi-target scenarios to solve the above-mentioned problems.

[0005] In a first aspect, embodiments of this application provide a target matching method for multi-target scenarios. The method includes: acquiring detection information of multiple human targets and multiple targets to be associated in a surveillance image; constructing a cost matrix based on the detection information; wherein the cost matrix is ​​used to characterize the matching cost between each human target and each target to be associated; the cost matrix integrates spatial distance constraints and prior human structure constraints, the prior human structure constraints being used to penalize candidate matches where the target to be associated is located outside the structural association region of the human target; solving the cost matrix to obtain the optimal matching result between the human target and the target to be associated; and filtering unreasonable matches based on the optimal matching result and a preset cost threshold to obtain the final matching result.

[0006] In the implementation of the above scheme, the target association problem in multi-person, multi-target scenarios is reconstructed into a globally optimal matching problem. By constructing a cost matrix and solving it, the globally optimal solution is obtained, avoiding the phenomenon of multiple matching of one target or mismatching due to proximity caused by local rule matching. On the other hand, the spatial distance constraint component quantifies the geometric proximity between the center point of the human body and the center point of the target to be associated, giving the matching process a clear spatial measurement benchmark and improving the accuracy of target spatial relationship judgment in complex scenarios. Furthermore, the human body structure prior constraint component encodes domain knowledge into structured constraints by penalizing candidate matches where the target to be associated is located outside the human body structure association area, reducing the probability of mismatching of non-associated targets. Finally, based on a preset cost threshold, the optimal matching results are screened for rationality, filtering out high-cost matching pairs and improving the accuracy and robustness of the final judgment.

[0007] In one implementation of the first aspect, constructing the cost matrix includes: for a human target-target pair, determining a spatial distance constraint component and a human structure prior constraint component of the cost matrix; wherein, the spatial distance constraint component includes a first distance constraint, which is used to quantify the distance between the center point of the human target and the center point of the target to be associated; the human structure prior constraint component includes a first penalty term; when the center point of the target to be associated is located inside the structural association region of the human target, the first penalty term is a first preset value; when the center point of the target to be associated is located outside the structural association region, the first penalty term is a second preset value; the first preset value is less than the second preset value; and fusing the spatial distance constraint component and the human structure prior constraint component to obtain the cost matrix.

[0008] In the implementation of the above scheme, the cost matrix construction process independently determines the spatial distance constraint component and the prior human structure constraint component for the human target-to-be-associated target pair. By weightedly fusing geometric proximity measurement with structured semantic constraints, the matching cost evaluation can simultaneously quantify the spatial location distance characteristics of candidate matching pairs and the degree of conformity with prior human structure knowledge. This provides a comprehensive quantitative evaluation basis for optimal matching solution that combines spatial information richness and domain knowledge discrimination. On the other hand, the prior human structure constraint component adopts an internal and external differentiated preset value design. When the center point of the target to be associated is located inside the structural association region of the human target, a lower first preset value is assigned, and when it is located outside, a higher second preset value is assigned. The clear difference in the size of the preset values ​​strengthens the conformity requirement of the matching result to the spatial constraints of the human structure, effectively suppresses erroneous matching between the target to be associated and human targets with inconsistent structural positions, reduces the probability of mismatch, and improves the accuracy of the matching result.

[0009] In one implementation of the first aspect, when a face target associated with the human body target is detected, determining the spatial distance constraint component and the prior human structure constraint component of the cost matrix includes: determining a second distance constraint of the spatial distance constraint component; wherein the second distance constraint is used to quantify the distance between the center point of the target to be associated and the center point of the face; determining a second penalty term of the prior human structure constraint component; wherein, when the center point of the target to be associated is located inside the structural association region of the face target, the second penalty term is a third preset value; when the center point of the target to be associated is located outside the structural association region of the face target, the second penalty term is a fourth preset value; the third preset value is less than the fourth preset value.

[0010] In the implementation of the above scheme, when a face target is detected, distance constraints and structural penalty terms related to the substructure are added, so that the cost matrix integrates the dual spatial constraints at the human body level and the substructure level, constructs a hierarchical structural prior constraint system, and improves the conformity of the matching results to the spatial semantics of human-object interaction behavior. On the other hand, the second penalty term adopts a differentiated design of third and fourth preset values. When the center point of the target to be associated is located inside the structural association region of the face target, a smaller third preset value is assigned, and when it is outside, a larger fourth preset value is assigned. This strengthens the spatial constraint of the substructure region on the target to be associated and reduces the probability of candidate matching that does not match the position of the target to be associated with the substructure.

[0011] In one implementation of the first aspect, fusing the spatial distance constraint component and the prior human structure constraint component to obtain the cost matrix includes: performing weighted fusion of the spatial distance constraint component and the prior human structure constraint component based on fusion weights to obtain the cost matrix; wherein the fusion weights are used to quantify the relative contributions of the spatial distance constraint component and the prior human structure constraint component in the cost matrix.

[0012] In the implementation of the above scheme, a fusion weight is introduced to weight and fuse the spatial distance constraint component and the human body structure prior constraint component, so that the cost matrix construction process has parameter adjustment capability, supports flexible configuration of the relative importance of each constraint item according to the characteristics of different monitoring scenarios, and realizes the adaptability of the matching strategy to scene changes. On the other hand, the fusion weight directly quantifies the relative contribution of spatial distance constraints and human body structure prior constraints in the cost matrix, and achieves a dynamic balance between geometric proximity and structured semantic constraints by adjusting the weight values, providing a configurable cost evaluation framework for optimal matching solution.

[0013] In one implementation of the first aspect, solving the cost matrix to obtain the optimal matching result between the human target and the target to be associated includes: performing column reduction and row reduction on the cost matrix so that each row and each column produces at least one zero element to obtain a bipartite graph model; wherein, the bipartite graph model includes a first node set representing multiple human targets and a second node set representing multiple targets to be associated; determining the minimum cost matching scheme that covers all nodes in the bipartite graph model through zero element coverage detection and matrix element iterative adjustment; extracting the maximum zero element matching set corresponding to the minimum cost matching scheme to obtain the optimal matching result with the minimum global total cost.

[0014] In the implementation of the above scheme, the local matching problem is transformed into a global optimization problem through a complete solution process of row and column reduction, zero element coverage detection and matrix element iterative adjustment. This ensures that the one-to-one optimal matching result with the minimum global total cost is obtained in complex scenarios with multiple people and multiple targets, avoiding the local optimum trap caused by greedy strategies. On the other hand, the bipartite graph model adopts an explicit modeling method in which the first node set represents the human target and the second node set represents the target to be associated, so that the matching process has a clear graph theory foundation. The matching cost distribution is dynamically optimized through the matrix element iterative adjustment mechanism, which improves the robustness of the optimal matching result in occluded and dense scenarios.

[0015] In one implementation of the first aspect, before solving the cost matrix, the method further includes: when the number of human targets is inconsistent with the number of targets to be associated, expanding the cost matrix into a square matrix using virtual nodes; wherein the matrix elements corresponding to the virtual nodes are filled with preset constant values; the preset constant values ​​are greater than preset reasonable matching costs.

[0016] In the implementation of the above scheme, the non-square cost matrix is ​​converted into a square matrix through the virtual node expansion strategy, so that the solving algorithm can uniformly handle the situation where the number of human targets is inconsistent with the number of targets to be associated, thus expanding the applicable scenarios of the target matching method for multi-target scenarios. On the other hand, the matrix elements corresponding to the virtual nodes are filled with a preset constant value greater than the preset reasonable matching cost to prevent the pairing of real targets and virtual nodes in the optimal matching result, avoid meaningless matching, and ensure the effectiveness of the matching scheme.

[0017] In one implementation of the first aspect, filtering unreasonable matches based on the optimal matching result and a preset cost threshold includes: if the target to be associated fails to match successfully and the minimum matching cost value obtainable by the target to be associated is higher than a first threshold, the target to be associated is determined as a false detection and filtered; if the minimum matching cost value corresponding to the target to be associated is between a second threshold and the first threshold, the target to be associated is marked as a suspected candidate; wherein the second threshold is less than the first threshold; and temporal consistency verification is performed on the suspected candidate based on a multi-frame image sequence, and if the suspected candidate continuously meets the matching conditions within a preset number of frames, it is confirmed that the target to be associated and the corresponding human target have target interaction behavior.

[0018] In the implementation of the above scheme, a hierarchical filtering mechanism is constructed using a first threshold and a second threshold to classify unmatched targets into two categories: false detections and suspected candidates. Differentiated processing strategies are implemented for different categories to achieve a balance between false detection suppression and boundary condition preservation. On the other hand, temporal consistency verification based on multi-frame image sequences is introduced for suspected candidates. The target to be associated is required to continuously meet the matching conditions within a preset number of frames to be determined as a target interaction behavior. The temporal dimension information is used to smooth the uncertainty of single-frame matching results, thereby improving the stability and robustness of the final judgment result. In one implementation of the first aspect, obtaining the final matching result includes: performing temporal consistency verification on the optimal matching result in a series of consecutive multi-frame image sequences; if the matching relationship in the optimal matching result continues to exist in the series of consecutive multi-frame image sequences, then confirming that the human target corresponding to the matching relationship has target interaction behavior with the target to be associated.

[0019] In the implementation of the above scheme, the optimal matching result is verified for temporal consistency in a series of consecutive multi-frame image sequences. The matching relationship must exist continuously in multiple frames to confirm the target interaction behavior. This makes the judgment result based on the stability of the time dimension and avoids misjudgment of interaction behavior caused by instantaneous mismatch in a single frame. On the other hand, the matching relationship is continuously verified by a series of consecutive multi-frame image sequences, which effectively filters out instantaneous false matches caused by occlusion, detection jitter or environmental interference, and improves the reliability and robustness of the final matching result.

[0020] In one implementation of the first aspect, the target to be associated is a smoke target.

[0021] In the implementation of the above scheme, the target matching method is applied to the smoke target recognition scenario. By encoding the spatial association between the smoke target and the human hand or mouth area through human structure prior constraints, the cost matrix can quantify the degree of unreasonableness of the smoke target being located inside or outside the human smoking area, thus achieving accurate matching between the smoke target and the smoker in complex scenarios with multiple people. On the other hand, considering the characteristics of the smoke target being small in scale and easily occluded, a global optimal matching strategy is adopted to replace the local distance rule, avoiding mismatch between multiple smoke targets and multiple human targets in dense crowds, and improving the accuracy and reliability of the behavior subject localization in the recognition of smoking behavior in non-smoking places.

[0022] Secondly, embodiments of this application provide a target matching system for multi-target scenarios, including a host computer and an image acquisition device communicatively connected to the host computer, wherein: The image acquisition device is used to acquire monitoring images and send the monitoring images to the host computer; The host computer is used to acquire detection information of multiple human targets and multiple targets to be associated in the monitoring image; construct a cost matrix based on the detection information; wherein the cost matrix is ​​used to characterize the matching cost between each human target and each target to be associated; the cost matrix integrates spatial distance constraints and prior human structure constraints, the prior human structure constraints are used to penalize candidate matches where the target to be associated is located outside the structural association region of the human target; solve the cost matrix to obtain the optimal matching result between the human target and the target to be associated; based on the optimal matching result and a preset cost threshold, filter unreasonable matches to obtain the final matching result.

[0023] Thirdly, embodiments of this application provide an electronic device, including: a processor, a memory, and a communication bus, wherein the processor and the memory communicate with each other through the communication bus; the memory stores computer program instructions that can be executed by the processor, and the computer program instructions are read and executed by the processor to perform the method provided in the first aspect or any possible implementation of the first aspect.

[0024] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when read and executed by a processor, perform the method provided in the first aspect or any possible implementation thereof.

[0025] Fifthly, embodiments of this application provide a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the method provided by the first aspect or any possible implementation of the first aspect.

[0026] Other features and advantages of this application will be set forth in the following description and will be apparent in part from the description or may be learned by practicing embodiments of this application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims and drawings. Attached Figure Description

[0027] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0028] Figure 1 A flowchart illustrating the target matching method for multi-target scenarios provided in this application embodiment; Figure 2 A flowchart illustrating a target matching method for a multi-target scenario in an application scenario provided in this application embodiment; Figure 3 This is a schematic diagram of the detection results in a certain application scenario provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of a target matching system for multi-target scenarios provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0029] The embodiments of the technical solution of this application will now be described in detail with reference to the accompanying drawings. These embodiments are only used to more clearly illustrate the technical solution of this application and are therefore merely examples, and should not be used to limit the scope of protection of this application.

[0030] In practical applications, human-object interaction behavior recognition typically relies on video images acquired by front-end camera devices, which are then processed by a back-end intelligent analysis system. When multiple people appear simultaneously in a monitored scene, or even when multiple people interact with the target to be associated at the same time, the system needs to accurately associate the target with the corresponding human body to determine which person is engaging in a specific interaction, thereby achieving accurate alarms, locating the responsible party, and subsequent management. However, in complex real-world scenarios, monitoring footage often suffers from issues such as dense crowds, severe obstruction, significant changes in viewing angle, and small scale of the target to be associated, making human-object matching a key technical challenge in interaction behavior recognition systems. Currently, the mainstream technical solutions for human-object interaction behavior recognition typically include the following methods: 1. Rule-based matching method: After detecting human targets and targets to be associated, simple spatial rules are used for matching. The main steps are as follows: (1) Associating the center point of the target to be associated with the nearest human target; (2) Determining whether the target to be associated falls into a specific part of the human body; (3) Associating the target by using a fixed distance threshold or area overlap ratio (IoU). This method is prone to target mismatch in multi-person scenarios. When there are multiple human targets and multiple targets to be associated in the picture, the matching method based on the nearest distance or threshold is prone to mismatching a target to be associated with a non-actual person, resulting in false positives or false negatives.

[0031] 2. Single-target greedy matching method: When multiple targets to be associated are detected, they are matched one by one according to the detection confidence or distance from small to large. Once a match is successful, it does not participate in subsequent matching. This method only focuses on the local optimal relationship and cannot guarantee the optimality of the matching result as a whole. This is especially evident when the number of targets is inconsistent or the spatial distribution is complex.

[0032] 3. End-to-end deep learning-based methods: These methods treat interactive behaviors as a whole category, directly classifying video segments through action recognition networks or temporal networks to output whether interactive behavior exists, without explicitly distinguishing the relationship between the human body and the target to be associated. This approach treats each image as a whole, classifying the input image without performing finer-grained analysis of human behavior within the image, and therefore cannot pinpoint whether a specific person engaged in interactive behavior.

[0033] In summary, existing solutions generally treat the matching problem between the human body and the target to be associated as a local rule-based judgment problem, rather than a global optimization matching problem. Without a unified mathematical model, matching strategies often rely on empirical thresholds or heuristic rules, making it difficult to simultaneously ensure accuracy, stability, and scalability in complex scenarios.

[0034] In view of this, embodiments of this application provide a target matching method for multi-target scenarios. This method reconstructs the target association problem in multi-person, multi-target scenarios into a globally optimal matching problem. By constructing a cost matrix and solving it, the globally optimal solution is obtained, avoiding the phenomenon of multiple matching of one target or mismatching due to proximity caused by local rule matching. On the other hand, the spatial distance constraint component quantifies the geometric proximity between the center point of the human body and the center point of the target to be associated, giving the matching process a clear spatial measurement benchmark and improving the accuracy of target spatial relationship judgment in complex scenarios. Furthermore, the human body structure prior constraint component penalizes candidate matches where the target to be associated is located outside the human body structure association area, encoding domain knowledge into structured constraints, reducing the probability of mismatching of non-associated targets. Finally, based on a preset cost threshold, the optimal matching results are rationally screened, filtering out high-cost matching pairs, thereby improving the accuracy and robustness of the final judgment.

[0035] Please see Figure 1 The illustrated flowchart illustrates a target matching method for multi-target scenarios provided in this application embodiment. This target matching method for multi-target scenarios can be applied to electronic devices, which may include physical devices such as servers, PCs, tablets, or smartphones, or virtual devices such as virtual machines or containers. The electronic device can be a single device, a combination of multiple devices, or a cluster of a large number of devices. The aforementioned target matching method for multi-target scenarios may include: Step S110: Obtain detection information of multiple human targets and multiple targets to be associated in the monitoring image.

[0036] The aforementioned surveillance images can originate from at least one image acquisition device deployed in the surveillance scene. These devices include, but are not limited to, network cameras, analog cameras, digital video recorders, or mobile smart terminals with image acquisition capabilities. The image acquisition devices can be deployed at preset locations within the surveillance area via fixed installation or pan-tilt-zoom (PTZ) control, with the acquisition angle covering the target monitoring range. Surveillance images can be acquired as real-time video streams or offline video files and transmitted via wired networks, wireless networks, or local storage media to the electronic device executing the aforementioned target matching method for multi-target scenarios. The raw image data output by the image acquisition device can be encoded and compressed to form a standard video format, including but not limited to H.264, H.265, or MJPEG encoding formats. The frame rate, resolution, and exposure parameters of the surveillance images can be dynamically configured and adjusted according to scene lighting conditions, target movement speed, or system processing requirements. Furthermore, in an edge computing architecture, surveillance images can be pre-processed by edge computing nodes deployed at the surveillance site before being transmitted to the electronic device executing the aforementioned target matching method for multi-target scenarios; in a cloud computing architecture, surveillance images can be directly uploaded to the electronic device executing the aforementioned target matching method for multi-target scenarios for centralized analysis and processing.

[0037] The aforementioned detection information includes a structured dataset extracted from surveillance images to describe the target's location, shape, and category attributes. Specifically, this includes: the bounding box coordinates for each target, where the bounding box defines the target's spatial occupancy in the image plane using the coordinates of the rectangle's vertices (usually the coordinates of the top-left and bottom-right corners or the center point coordinates plus width and height parameters). The detection information may also include a target category identifier to distinguish between different object categories such as human targets and targets to be associated. Furthermore, it may include a detection confidence score, characterizing the model's certainty about the target's existence. Technically, the detection information can be organized using array, list, or hash table data structures, with each target instance corresponding to a data record. In another implementation, the detection information may also include a set of pixel-level coordinates of the target segmentation mask, a sequence of keypoint coordinates, or a deep feature vector to enhance the feature representation capability in subsequent matching stages.

[0038] The above step S110 can obtain detection information through at least one of the following methods; The first method: using a deep learning object detection model to obtain the target; Deep learning object detection models can run on cloud servers, edge computing nodes, or local processing units. The model receives surveillance images as input, extracts multi-scale features through convolutional neural networks, and performs classification and regression tasks, outputting the bounding box coordinates, detection confidence, and category label for each object. In a distributed architecture, edge nodes can perform preliminary detection and transmit the results to the cloud for aggregation processing; in a centralized architecture, the raw video stream is directly uploaded to the cloud for unified detection. Deep learning object detection models can employ two-stage detectors (such as Faster R-CNN), single-stage detectors (such as YOLO, SSD), or Transformer-based detectors (such as DETR) architectures. Detection information is organized using structured data formats (such as JSON, XML, or binary protocols).

[0039] The second method: using a time-based sequence processing mechanism to obtain it; A time-based sequence processing mechanism can be used to extract single-frame images or image sequence blocks from video streams or offline video files at preset time intervals, and target detection can be performed independently for each frame or sequence block. The detection process can be implemented by calling computer vision library functions or by loading the weight file of a pre-trained neural network for forward inference. Multi-frame detection results can carry timestamp information to achieve temporal alignment, and a variable-length data structure is used to store detection information when the number of targets changes dynamically. In batch processing mode, consecutive frames of images can be grouped into batches for input to the model to improve computational efficiency, and the output results are doubly identified by frame number and target index.

[0040] The third method: using a multi-task joint detection framework to obtain; A multi-task joint detection framework is used to simultaneously output detection information for human targets and targets to be associated using a single model. The backbone network of the model shares a feature extraction layer, and after feature fusion in the neck network, each target category is output separately in different detection heads. The multi-task joint detection framework can reduce redundant computational overhead and ensure the time synchronization of multi-class detection. In addition to bounding boxes, the detection information can also include target mask pixel sets, keypoint coordinates, or deep feature embedding vectors to improve the feature discrimination ability in the subsequent matching stage. During the model deployment stage, model compression techniques (such as quantization and pruning) can be used to optimize inference speed, and hardware acceleration units (such as GPUs, TPUs, or NPUs) can be used to improve processing throughput.

[0041] Step S120: Based on the detection information, construct a cost matrix; wherein, the cost matrix is ​​used to characterize the matching cost between each human target and each target to be associated; the cost matrix integrates spatial distance constraints and human structure prior constraints, and the human structure prior constraints are used to penalize candidate matches where the target to be associated is located outside the structural association region of the human target.

[0042] The aforementioned cost matrix is ​​a two-dimensional array data structure based on the real number field. Its row dimension corresponds to a set of multiple human targets, and its column dimension corresponds to a set of multiple targets to be associated. Each element in the matrix stores a quantitative evaluation value of the matching relationship between a specific human target and a specific target to be associated. This matrix integrates spatial distance constraints and prior constraints of human structure into a single numerical index through mathematical modeling, which is used to characterize the comprehensive matching cost of candidate matching pairs in terms of both geometric spatial proximity and structural semantic rationality. Among them, the spatial distance constraint component reflects the Euclidean distance or metric distance between the center point of the human body and the center point of the target to be associated in the image coordinate system, while the prior constraint component of human structure imposes an additional cost increment on candidate matches that violate the spatial association rules of human structure through a penalty mechanism. When the center point of the target to be associated is located outside the structural association region of the human target, the prior constraint of human structure penalizes the candidate match by increasing the matching cost value, so that the cost matrix not only contains pure geometric information, but also encodes domain prior knowledge, providing a mathematical representation foundation with both spatial metric and semantic discriminative capabilities for subsequent global optimization solutions.

[0043] It is understandable that in complex monitoring scenarios involving multiple people and multiple targets, relying solely on geometric location information is insufficient to fully describe the semantic rationality of human-target interaction. Therefore, the target matching method for multi-target scenarios provided in this application employs a fusion mechanism of spatial distance constraints and prior constraints on human structure. Specifically, the spatial distance constraint quantifies the Euclidean distance between the center point of the human target and the center point of the target to be associated, providing a basic geometric proximity metric for the matching process. This ensures the feasibility of candidate matching pairs in physical space and effectively eliminates irrelevant target combinations that are too far apart. The prior constraints on human structure penalize candidate matches where the target to be associated is located outside the human structure association region. This encodes domain knowledge into a computable cost increment, ensuring that the cost matrix reflects not only geometric distance but also the semantic rules of human spatial structure. This distinguishes between valid matches that conform to the spatial pattern of interaction behavior and invalid matches that violate structural common sense, improving the discrimination ability and robustness of the matching results in complex scenarios.

[0044] Optionally, step S120 above, which constructs a cost matrix, includes: for the human target-target pair, determining the spatial distance constraint component and the human structure prior constraint component of the cost matrix; wherein, the spatial distance constraint component includes a first distance constraint, which is used to quantify the distance between the center point of the human body and the center point of the target to be associated; the human structure prior constraint component includes a first penalty term; when the center point of the target to be associated is located inside the structural association region of the human target, the first penalty term is a first preset value; when the center point of the target to be associated is located outside the structural association region, the first penalty term is a second preset value; the first preset value is less than the second preset value; and fusing the spatial distance constraint component and the human structure prior constraint component to obtain the cost matrix.

[0045] The aforementioned first distance constraint is a geometric metric based on Euclidean distance. It quantifies the degree of matching in terms of positional proximity between candidate matching pairs by calculating the spatial distance in the image plane between the center coordinates of the human target's bounding box and the center coordinates of the target's bounding box to be associated. The first distance constraint converts the difference in center coordinates between two targets into a scalar distance value, which serves as the basic spatial component in the cost matrix for evaluating the reasonableness of the match. A smaller distance indicates a closer spatial relationship between the candidate matching pairs and a higher probability of matching.

[0046] The first penalty term mentioned above is a discrete penalty mechanism in the prior constraint component of human body structure, used to encode the spatial attribution rule of the target to be associated relative to the associated region of human body structure. The first penalty term assigns different preset costs based on whether the center point of the target to be associated is located within the associated region of the human body structure: a smaller first preset value is assigned when the center point is within the region, indicating compliance with prior knowledge of human body structure; a larger second preset value is assigned when the center point is outside the region, imposing additional cost penalties to suppress candidate matches that violate spatial semantic rules. The difference between the first and second preset values ​​strengthens the requirement for the matching results to conform to the human body spatial structure, effectively filtering out matching pairs consistent with the spatial patterns of human interaction behavior.

[0047] The aforementioned structural association region refers to a spatial constraint range defined based on human anatomical features or specific interaction behavior patterns, used to quantify the spatial semantic rationality between the target to be associated and human body parts. This region is constructed using the geometric center or bounding box of the human target as a reference, through preset spatial offsets and region size parameters, representing the set of reasonable spatial locations where the target to be associated may appear in a specific interaction behavior. In a smoking detection scenario, the structural association region corresponds to the human hand or mouth area; in a flame detection scenario, it corresponds to the area where the human is holding an object; and in a hazardous material detection scenario, it corresponds to the area where the human is carrying something.

[0048] The aforementioned structural association regions can be determined using at least one of the following methods: The first method is the bounding box scaling method, which generates candidate regions by scaling down the human target detection box proportionally, with the scaling factor set according to prior knowledge of human body parts; the second method is the target detection method, which directly outputs the bounding boxes of specific human body parts (such as the head and hands) as structural association regions through an independent target detection model; the third method is the keypoint mapping method, which uses a human pose estimation model to obtain keypoint coordinates and expands a fixed radius region centered on the keypoints or adaptively generates regions based on human body proportions; the fourth method is the semantic segmentation method, which generates pixel-level semantic masks through a human body part segmentation model and uses the set of pixels belonging to a specific part as the accurate structural association region.

[0049] The aforementioned first preset value can be set to zero or a constant value close to zero to completely eliminate or minimize the penalty for candidate matches that conform to prior knowledge of human structure. This indicates that such matches are reasonable at the spatial semantic level and should not incur additional costs due to structural constraints. This preset value can be determined through parameter search on an offline validation set, taking the minimum value while ensuring that effective matches are not suppressed; alternatively, it can be set based on the distance equivalence principle, mapping it to a numerical scale equivalent to typical distance constraint terms to ensure that structural constraint components and spatial distance constraint components are comparable in the cost matrix.

[0050] The second preset value can be set to a positive real number greater than the first preset value. Its magnitude can be configured according to the geometric dimensions of the structurally associated region and the spatial scale of the monitoring scene. It can typically be set to a fixed multiple of the diagonal length of the structurally associated region, so that the penalty intensity and the degree of spatial deviation form a reasonable mapping relationship. This preset value can be determined through experimental verification, gradually increasing the value in multiple test scenarios until the matching pairs that violate the structural prior are effectively suppressed, while avoiding missed detections due to excessive penalty; alternatively, an adaptive adjustment mechanism can be adopted to dynamically adjust the penalty intensity according to the scene complexity, achieving a balance between mismatch suppression and matching recall.

[0051] In the above scheme, the cost matrix construction process independently determines the spatial distance constraint component and the prior human structure constraint component for the human target-to-be-associated target pair. By weightedly fusing geometric proximity measurement with structured semantic constraints, the matching cost evaluation can simultaneously quantify the spatial location distance characteristics of candidate matching pairs and the degree of conformity with prior human structure knowledge. This provides a comprehensive quantitative evaluation basis for optimal matching solution that combines spatial information richness and domain knowledge discrimination. On the other hand, the prior human structure constraint component adopts an internal and external differentiated preset value design. When the center point of the target to be associated is located inside the structural association region of the human target, a lower first preset value is assigned, and when it is located outside, a higher second preset value is assigned. The clear difference in the size of the preset values ​​strengthens the conformity requirement of the matching result to the spatial constraints of the human structure, effectively suppressing erroneous matching between the target to be associated and human targets with inconsistent structural positions, reducing the probability of mismatch and improving the accuracy of the matching result.

[0052] Optionally, when a face target associated with a human target is detected, the determination of the spatial distance constraint component and the prior human structure constraint component of the cost matrix includes: determining a second distance constraint for the spatial distance constraint component; wherein the second distance constraint is used to quantify the distance between the center point of the target to be associated and the center point of the face; determining a second penalty term for the prior human structure constraint component; wherein, when the center point of the target to be associated is located inside the structural association region of the face target, the second penalty term is a third preset value; when the center point of the target to be associated is located outside the structural association region of the face target, the second penalty term is a fourth preset value; the third preset value is less than the fourth preset value.

[0053] It is understandable that in multi-person interaction scenarios, the face region, as a key substructure of the human body, can provide a more refined spatial semantic reference than the overall human body bounding box. The purpose of adding a second distance constraint and a second penalty term when detecting a face target in the above scheme is that the distance metric between the center point of the face and the center point of the target to be associated can capture more granular spatial interaction patterns. Especially when human-object interactions are highly concentrated in the facial region, face-level distance constraints can effectively distinguish different interaction intentions in similar spatial locations. Simultaneously, the second penalty term constructs a hierarchical structural prior constraint system by determining whether the target to be associated is located within the face structure association region. This allows the cost matrix to simultaneously encode both human-level and face-level spatial rules, strengthening the matching results' preference for candidate pairs that conform to the spatial distribution characteristics of the face region, reducing the probability of mismatches due to conflicts between the target to be associated and the face position, and improving the spatial semantic consistency of the matching results in complex poses or densely occluded scenarios.

[0054] The aforementioned second distance constraint is a geometric metric based on Euclidean distance. It quantifies the degree of matching between candidate matching pairs in terms of facial region proximity by calculating the spatial distance between the center coordinates of the bounding box of the target to be associated and the center coordinates of the bounding box of the face target. This constraint converts the spatial coordinate difference between the target to be associated and the face target into a scalar distance value, serving as a supplementary spatial component in the cost matrix for evaluating the likelihood of facial region interaction in candidate matches. A smaller distance indicates a closer spatial relationship between the candidate matching pairs in the facial region, and a higher probability of face-related interactions.

[0055] The second penalty term mentioned above is a discrete penalty mechanism for the face region within the prior constraint component of human body structure. It is used to encode the spatial attribution rule of the target to be associated relative to the face structure association region. This penalty term assigns different preset costs based on whether the center point of the target to be associated is located inside the structure association region of the face target: a smaller third preset value is assigned when the center point of the target to be associated is within the region, indicating that it conforms to the prior knowledge of the face region structure; a larger fourth preset value is assigned when it is outside the region, imposing an additional cost penalty to suppress candidate matches that violate the face spatial semantic rules. The differentiated design of the third and fourth preset values ​​strengthens the conformity of the matching results to the face region spatial constraints, effectively filtering out matching pairs consistent with the spatial patterns of face interaction behavior.

[0056] Similar to the setting methods of the first and second preset values, the third preset value can be set to zero or a constant close to zero. This is used to eliminate or minimize the penalty for candidate matches that conform to prior knowledge of face structure, indicating that such matches are reasonable at the spatial semantic level of the face region. The third preset value can be determined through parameter search on an offline validation set, taking the minimum value while ensuring that effective matches related to the face region are not suppressed. Alternatively, it can be mapped to a numerical scale equivalent to typical spatial distance constraints based on the distance equivalence principle, ensuring that the contribution weights of face-level structural constraint components and other constraint components in the cost matrix are comparable. The fourth preset value can be set to a positive integer greater than the third preset value. Its numerical magnitude needs to be configured according to the geometric dimensions of the face structure-related region and the spatial scale of the monitoring scene. It can usually be set to a fixed multiple of the diagonal length of the face structure-related region, so that the face region applies sufficient penalty intensity to deviations from the target. The fourth preset value can be determined through experimental verification. The value can be gradually increased in multiple test scenarios until the matching pairs that violate the prior knowledge of the face region structure are effectively suppressed. At the same time, it avoids excessive punishment that may lead to missed detection of effective interaction behaviors related to the face. An adaptive adjustment mechanism can also be used to dynamically adjust the punishment intensity according to the complexity of the scenario.

[0057] The above scheme adds distance constraints and structural penalty terms regarding substructures when detecting a face target, enabling the cost matrix to integrate dual spatial constraints at the human body level and substructure level, constructing a hierarchical structural prior constraint system, and improving the consistency of the matching results with the spatial semantics of human-object interaction behavior. On the other hand, the second penalty term adopts a differentiated design of third and fourth preset values. When the center point of the target to be associated is located inside the structural association region of the face target, a smaller third preset value is assigned, and when it is outside, a larger fourth preset value is assigned. This strengthens the spatial constraint of the substructure region on the target to be associated and reduces the probability of candidate matching that does not match the position of the target to be associated with the substructure.

[0058] Optionally, the above-mentioned fusion of spatial distance constraint components and human structure prior constraint components to obtain the cost matrix includes: weighting and fusing the spatial distance constraint components and human structure prior constraint components based on fusion weights to obtain the cost matrix; wherein, the fusion weights are used to quantify the relative contributions of the spatial distance constraint components and human structure prior constraint components in the cost matrix.

[0059] When setting the fusion weights of various constraints in the cost matrix, the following factors can be considered: (1) Geometric characteristics and spatial scale distribution of the monitoring scene: In open areas with a wide scene range and large target spacing, a larger spatial distance constraint weight can be configured to ensure that geometric proximity dominates the matching decision; in closed spaces with a limited scene range and densely distributed targets, the weight of the prior constraint of human structure can be increased to strengthen the discriminative power of spatial semantic rules and avoid mismatch due to close distance. (2) Target distribution density and occlusion degree: In complex scenes with dense human targets and severe mutual occlusion, the detection box positioning error increases and the spatial distance discrimination decreases. At this time, the weight of the prior constraint of structure can be increased to make up for the lack of geometric information by using the semantic information of the human structure associated region; in ideal scenes with sparse targets and less occlusion, the weight of the structural constraint can be appropriately reduced to avoid over-penalizing effective matching. (3) Accuracy and positioning error characteristics of the detection model: When the confidence of the bounding box coordinates output by the target detection model is low or the positioning jitter is obvious, the structural prior constraint weight can be increased as a robust compensation mechanism to suppress the impact of detection noise on the matching results; when the detection model has high-precision positioning capability, the spatial distance constraint weight can be increased accordingly to make full use of accurate geometric information. (4) Performance index requirements of the application task: For security supervision scenarios with high recall priority, the structural constraint penalty intensity can be appropriately reduced to avoid the risk of missed detection; for compliance monitoring scenarios with strict accuracy requirements, the structural prior constraint weight can be enhanced to minimize false alarms.

[0060] The aforementioned weighted fusion mechanism, by introducing configurable weight parameters, dynamically adjusts the relative contributions of different constraint components in the cost matrix, enabling the matching strategy to adapt to the specific needs of different monitoring scenarios. In densely occluded scenarios, the weight of structural prior constraints can be increased to strengthen the binding force of spatial semantic rules; in open scenarios, the weight of spatial distance constraints can be increased to prioritize geometric proximity. The weighted fusion mechanism provides scenario-adaptive capabilities for system deployment, determines the optimal weight configuration through an offline validation set, balances the mutual influence between geometric metrics and semantic constraints, avoids matching bias caused by excessive dominance of a single constraint, and improves the robustness of cost evaluation to diverse human-object interaction patterns.

[0061] The above scheme introduces fusion weights to weight and fuse spatial distance constraint components and human structure prior constraint components, enabling the cost matrix construction process to have parameterized adjustment capabilities. This allows for flexible configuration of the relative importance of each constraint item based on the characteristics of different monitoring scenarios, achieving adaptability of the matching strategy to scenario changes. On the other hand, the fusion weights directly quantify the relative contributions of spatial distance constraints and human structure prior constraints in the cost matrix. By adjusting the weight values, a dynamic balance between geometric proximity and structured semantic constraints is achieved, providing a configurable cost evaluation framework for optimal matching solutions.

[0062] Step S130: Solve the cost matrix to obtain the optimal matching result between the human target and the target to be associated.

[0063] The aforementioned cost matrix calculation is a mathematical process that formalizes the multi-objective matching problem into a combinatorial optimization problem and obtains the globally optimal solution. Given the cost matrix, the goal is to find a set of matching relationships that assigns each human target to each target to be associated, ensuring that each human target matches at most one target to be associated, and each target to be associated matches at most one human target, thus minimizing the total matching cost of all matching pairs globally. This process differs from local greedy matching strategies; its optimization scope covers all candidate matching combinations, evaluating the total cost of the matching scheme as a whole and avoiding local optima traps. The solution output is a set of matching pair indices, where each matching pair indicates the optimal association between a specific human target and a specific target to be associated. Unmatched targets are considered unassociated objects. This solution process balances computational complexity with the optimality of the solution, using a polynomial-time optimization algorithm to ensure a globally optimal matching scheme in multi-objective scenarios with inconsistent target numbers or complex spatial distributions, providing a mathematically optimal target association foundation for subsequent behavior determination.

[0064] Step S130 above can employ the Hungarian algorithm to solve for the cost matrix. The Hungarian algorithm is a classic combinatorial optimization method for solving the minimum weight matching problem in bipartite graphs, achieving a globally optimal matching scheme within polynomial time complexity. The Hungarian algorithm abstracts the human target set and the set of targets to be associated as two vertex sets in a bipartite graph. The elements in the cost matrix correspond to the weights of the edges in the bipartite graph. Through matrix transformation and optimization iteration, it finds the perfect match that minimizes the total weight. This algorithm guarantees that each human target matches at most one target to be associated, and each target to be associated matches at most one human target, avoiding matching conflicts. The matching result output by the Hungarian algorithm is a one-to-one optimal allocation scheme, minimizing the total cost of all successful matches globally. This effectively handles multi-target scenarios with inconsistent target numbers or complex spatial distributions, providing a mathematically optimal target association basis for behavior recognition.

[0065] Optionally, step S130 may include: performing column and row reduction on the cost matrix to generate at least one zero element in each row and column, thereby obtaining a bipartite graph model; wherein the bipartite graph model includes a first set of nodes representing multiple human targets and a second set of nodes representing multiple targets to be associated; determining the minimum cost matching scheme that covers all nodes in the bipartite graph model through zero-element coverage detection and matrix element iterative adjustment; extracting the maximum matching set of zero elements corresponding to the minimum cost matching scheme to obtain the optimal matching result with the minimum global total cost. An example of this implementation is: (1) Perform row and column reduction operations on the cost matrix, subtracting the smallest element of each row and then the smallest element of each column, so that each row and column of the matrix produces at least one zero element. This transformation does not change the optimal matching set, but only adjusts the cost scale. Based on the reduced matrix, construct a bipartite graph model, abstracting multiple human targets as a first node set and multiple targets to be associated as a second node set. The zero element position corresponds to the candidate matching edge between the two node sets, forming a graph structure representation that can be used for subsequent search.

[0066] (2) Determine whether the current set of zero elements contains a perfect match by checking the zero element coverage, i.e., whether all zero elements can be covered by the minimum number of lines and the coverage number is equal to the matrix order. If this condition is not met, iteratively adjust the matrix elements: determine the minimum value among the uncovered elements, subtract the minimum value from all uncovered elements, and add the minimum value to the elements at the intersection of rows and columns. This operation generates new zero elements while keeping the original zero elements unchanged. Repeat the coverage check and matrix adjustment process until there is a perfect match covering all nodes in the set of zero elements. At this time, the position of the zero element in the matrix corresponds to a set of candidate matching pairs that do not conflict and have the minimum total cost.

[0067] (3) Extract the set of matching with the largest cardinality from the zero elements of the final matrix, that is, find the combination of independent zero elements that covers the most nodes. This combination corresponds to the optimal one-to-one matching relationship between the human target and the target to be associated. The extraction process ensures that each human target matches at most one target to be associated, and each target to be associated matches at most one human target. The resulting matching scheme has the minimum global total cost. Targets that do not participate in the matching are regarded as unassociated objects, thus completing the solution from the cost matrix to the optimal matching result.

[0068] The above scheme transforms the local matching problem into a global optimization problem through a complete solution process of row and column reduction, zero-element coverage detection, and matrix element iterative adjustment. This ensures that a one-to-one optimal matching result with the minimum global total cost is obtained in complex scenarios with multiple people and multiple targets, avoiding the local optimum trap caused by greedy strategies. On the other hand, the bipartite graph model adopts an explicit modeling method in which the first node set represents the human target and the second node set represents the target to be associated, giving the matching process a clear graph theory foundation. The matching cost distribution is dynamically optimized through the matrix element iterative adjustment mechanism, improving the robustness of the optimal matching result in occluded and dense scenarios.

[0069] Optionally, before solving the cost matrix, the above-mentioned target matching method for multi-target scenarios may further include: when the number of human targets is inconsistent with the number of targets to be associated, expanding the cost matrix into a square matrix using virtual nodes; wherein the matrix elements corresponding to the virtual nodes are filled with preset constant values; the preset constant values ​​are greater than preset reasonable matching costs.

[0070] The Hungarian algorithm, as a standard method for solving the minimum weight matching problem in bipartite graphs, requires the cost matrix to be a square matrix, meaning the number of rows and columns must be equal. When the number of human targets differs from the number of targets to be associated, the cost matrix becomes non-square. Directly applying the Hungarian algorithm leads to a dimension mismatch problem, making it impossible to complete the matrix reduction and optimization iteration operations. By introducing virtual nodes to expand the square matrix, the dimensions of the fewer targets are supplemented to be equal to the dimensions of the more numerous targets by adding virtual rows or columns. This ensures that the cost matrix meets the dimensional constraints of the algorithm input, allowing the optimization process to completely traverse all candidate matching relationships between real and virtual targets, guaranteeing the mathematical feasibility and algorithmic integrity of the solution process.

[0071] The preset constant value filling the matrix elements corresponding to virtual nodes can be set to an upper bound greater than the preset reasonable matching cost range. This preset constant value prevents the formation of lower-cost matching pairs between real targets and virtual nodes. The preset reasonable matching cost can be pre-calculated based on the spatial scale of the monitoring scene, the size distribution of the target detection box, and the distance constraint weight parameters. It is typically based on the theoretical maximum value of all elements in the cost matrix or the maximum matching cost actually observed. The preset constant value is increased by a fixed multiple or a fixed offset to ensure that its magnitude is significantly higher than any possible real matching cost. This allows the Hungarian algorithm to prioritize matching combinations between real targets during optimization, avoiding the incorrect assignment of real targets to virtual nodes representing no matching meaning, and maintaining the physical meaning and effectiveness of the optimal matching result.

[0072] The above scheme transforms the non-square cost matrix into a square matrix through a virtual node expansion strategy, enabling the solution algorithm to uniformly handle situations where the number of human targets is inconsistent with the number of targets to be associated, thus expanding the applicable scenarios of the target matching method for multi-target scenarios. On the other hand, the matrix elements corresponding to the virtual nodes are filled with preset constant values ​​greater than the preset reasonable matching cost to prevent real targets from being paired with virtual nodes in the optimal matching result, avoid meaningless matching, and ensure the effectiveness of the matching scheme.

[0073] Step S140: Based on the optimal matching result and the preset cost threshold, filter out unreasonable matches and obtain the final matching result.

[0074] The aforementioned preset cost threshold is a boundary parameter set to determine the rationality of a match. It is used to compare the cost value with that of the optimal matching result. When the cost value exceeds this threshold, the matching pair is judged as an unreasonable match and filtered out. This threshold can be configured as a single value or a set of segmented values, forming a hard decision boundary in the cost evaluation system. This divides the matching result space into valid and invalid matching regions, enabling automatic removal of low-quality matches and ensuring the overall credibility and accuracy of the final matching result set.

[0075] When setting the above preset cost threshold, the following factors can be considered: (1) Geometric distance scale of the monitoring scene: In scenes with a wide range and large target spacing, a higher threshold can be set to avoid over-filtering of effective matches. In scenes with a compact range and dense targets, the threshold can be reduced to improve the sensitivity to unreasonable spatial matching. (2) Positioning error and jitter characteristics of the target detection model: When there is a large uncertainty in the coordinates of the detection box, the threshold should be appropriately increased to ensure the tolerance to detection noise. (3) Performance requirements of the application task for recall and precision: In security supervision scenarios where recall is prioritized, the threshold can be relaxed to reduce missed detections. In compliance monitoring scenarios where precision is prioritized, the threshold can be tightened to minimize false alarms.

[0076] Optionally, step S140 above filters unreasonable matches based on the optimal matching result and a preset cost threshold, including: if the target to be associated fails to match successfully, and the minimum matching cost value that the target to be associated can obtain is higher than the first threshold, the target to be associated is judged as a false detection and filtered; if the minimum matching cost value corresponding to the target to be associated is between the second threshold and the first threshold, the target to be associated is marked as a suspected candidate; wherein, the second threshold is less than the first threshold; the temporal consistency verification of the suspected candidate is performed based on a multi-frame image sequence, and if the suspected candidate continuously meets the matching conditions within a preset number of frames, it is confirmed that there is target interaction behavior between the target to be associated and the corresponding human target.

[0077] When the target to be associated fails to find a match, it indicates that the target cannot find a pair that satisfies the minimum cost criterion during the global optimal matching process. In other words, the matching cost between the target and all human targets does not meet the global optimality requirement. If the minimum matching cost achievable by the target is still higher than the first threshold, it means that even the weighted sum of the spatial distance constraint and structural prior constraint between the target and the human candidate with the minimum cost is still too large. This implies that the target's position in the image space has a significant spatial deviation or semantic conflict with the structural association regions of all human targets. In this case, the target to be associated is highly likely to be a false alarm target generated by the detection model, and its appearance in the image does not conform to the spatial distribution pattern that human-object interaction behavior should have. Identifying and filtering such targets as false detections can effectively suppress the influence of false targets caused by environmental noise, lighting interference, or oversensitivity of the detection model on subsequent behavior judgments, thereby improving the accuracy and reliability of the final interaction behavior recognition results.

[0078] When the minimum matching cost of the target to be associated falls between the second threshold and the first threshold, it indicates that although the matching cost between the target and the human target is higher than the strict acceptance standard, it has not yet reached the level of complete rejection, belonging to the transitional zone between spatial rationality and irrationality. Targets between the first and second thresholds may have high matching costs due to detection and positioning deviations, slight occlusion, or abnormal human posture, rather than being completely isolated false alarm targets. Therefore, single-frame judgment is prone to misjudgment or missed judgment. Marking such targets as suspected candidates and introducing a temporal consistency verification mechanism can accumulate evidence in the time dimension to improve the confidence of the judgment. By requiring the target to continuously meet the matching conditions within a preset number of frames, it effectively distinguishes between accidental matching caused by instantaneous interference and real continuous human-object interaction behavior, reducing the uncertainty of single-frame decision. Temporal consistency verification tracks and statistically analyzes the matching status of suspected candidates in consecutive frames. When a candidate stably maintains a matching relationship within a time period exceeding the preset frame threshold, it is judged as target interaction behavior. The temporal consistency verification mechanism utilizes temporal information to smooth out random fluctuations in single-frame matching results, filtering out momentary false matches caused by detection jitter, sudden changes in illumination, or brief occlusions, ensuring that the final judgment is based on temporal stability. Simultaneously, for targets within the boundary cost range, temporal verification provides additional decision-making support, avoiding the missed detection of genuine but poorly localized interactions due to overly strict threshold settings. This achieves a balance between false detection suppression and recall, improving the robustness and reliability of the method in complex dynamic scenarios.

[0079] The above scheme constructs a hierarchical filtering mechanism through a first threshold and a second threshold, classifying unmatched targets to be associated into two categories: false detections and suspected candidates. Differentiated processing strategies are implemented for different categories to achieve a balance between false detection suppression and boundary condition preservation. On the other hand, temporal consistency verification based on multi-frame image sequences is introduced for suspected candidates. The target to be associated is required to continuously meet the matching conditions within a preset number of frames before it can be determined as a target interaction behavior. The temporal dimension information is used to smooth the uncertainty of the single-frame matching result and improve the stability and robustness of the final judgment result.

[0080] Optionally, step S140 above, which obtains the final matching result, includes: performing temporal consistency verification on the optimal matching result in a series of consecutive frames of images; if the matching relationship in the optimal matching result continues to exist in the series of consecutive frames of images, then it is confirmed that the human target corresponding to the matching relationship and the target to be associated have target interaction behavior.

[0081] A temporal consistency verification mechanism is introduced for the optimal matching result. By continuously tracking the stability of the matching relationship in a series of consecutive image frames, the matching pair must remain continuously present for a period exceeding a preset frame threshold in order to finally confirm the target interaction behavior. The temporal consistency mechanism incorporates accumulated evidence from the time dimension into the judgment process, uses multi-frame information to smooth the random fluctuations of single-frame matching results, effectively filters false matches caused by instantaneous interference, and ensures that the judgment result is based on spatiotemporal dual verification, significantly improving the reliability and robustness of the final interaction behavior recognition.

[0082] The above scheme verifies the temporal consistency of the optimal matching result in a series of consecutive image frames. It requires that the matching relationship persists across multiple frames to confirm the target interaction behavior, thus establishing the judgment result on the basis of temporal stability and avoiding misjudgment of interaction behavior caused by momentary false matching in a single frame. On the other hand, by using a series of consecutive image frames to continuously verify the matching relationship, it effectively filters out momentary false matching caused by occlusion, detection jitter, or environmental interference, thereby improving the reliability and robustness of the final matching result.

[0083] The targets to be associated mentioned above can be one or more of the following forms: The first type: The target to be associated is a smoke target; When the target to be associated is a smoke target, the above-mentioned target matching method for multi-target scenarios can be applied to the smoking behavior recognition scenario. In this scenario, the human body structure prior constraint is constructed based on the spatial association between the smoke target and the human hand area or mouth area. The reasonableness of the matching of the smoke target located inside or outside the human smoking area is evaluated by the cost matrix, so as to achieve accurate matching and positioning of the smoke target and the smoker in the no-smoking areas such as factory workshops, gas stations, and hospitals.

[0084] The above scheme applies the target matching method to the smoke target recognition scenario. By encoding the spatial association between the smoke target and the hand or mouth region of the human body through prior constraints on human structure, the cost matrix can quantify the degree of unreasonableness of the smoke target being located inside or outside the human body's smoking area, thus achieving accurate matching between the smoke target and the smoker in complex scenarios with multiple people. On the other hand, considering the characteristics of smoke targets being small in scale and easily occluded, a global optimal matching strategy is adopted to replace the local distance rule, avoiding mismatches between multiple smoke targets and multiple human targets in dense crowds, and improving the accuracy and reliability of the behavior subject localization in the recognition of smoking behavior in non-smoking places.

[0085] The second type: The target to be associated is a fire target; When the target to be associated is a flame target, the above-mentioned target matching method for multi-target scenarios can be applied to fire risk early warning application scenarios. In this scenario, the prior constraint of human body structure is constructed based on the spatial relationship between the flame target and the human hand area or the area where the object is held. By judging whether the center point of the flame is located inside the human body bounding box or the hand structure associated area, the arson or illegal use of fire behavior can be identified, and the flame target and the responsible personnel in the storage area and flammable material storage area can be accurately associated.

[0086] The third type: The target to be associated is a dangerous goods target; When the target to be associated is a hazardous material, the target matching method for multi-target scenarios described above can be applied to industrial safety supervision scenarios. In this scenario, the prior constraints of human body structure are constructed based on the spatial association between the hazardous material target and the human hand area or carrying area. The spatial attachment relationship between hazardous materials and personnel is evaluated through the cost matrix, and behaviors such as illegally carrying hazardous materials into the production area, intrusion into the dangerous area, or failure to use hazardous materials in accordance with the prescribed operating procedures are identified, so as to achieve accurate correspondence between hazardous materials and relevant personnel.

[0087] The fourth type: The target to be associated is a specific tool or equipment target; When the target to be associated is a specific tool or equipment target, the above-mentioned target matching method for multi-target scenarios can be applied to the production process compliance monitoring scenario. In this scenario, the prior constraint of human body structure is constructed based on the spatial relationship between the tool target and the human hand area. By judging whether the center point of the tool is within the operable range of the hand, the system monitors whether the operator uses the specified tool as required, whether the operation steps are correctly executed, or whether there is any abuse of the tool, thereby achieving accurate matching between the tool target and the operator and determination of behavioral compliance.

[0088] It is understandable that when applying target matching methods to different scenarios such as smoking detection, fire early warning, hazardous materials supervision, or tool compliance monitoring, the method parameters can be configured differently based on the spatial scale characteristics, target distribution density, detection model accuracy, and task performance requirements of each scenario. Parameter configuration objects include the fusion weights of spatial distance constraints and prior human structure constraints, the preset penalty values ​​for human-level and face-level constraint terms, the constant filling values ​​of virtual nodes during cost matrix expansion, the first and second cost thresholds for the filtering stage, and the preset frame number threshold for temporal consistency verification. The configuration process can perform parameter search and performance evaluation based on offline validation sets. For example, considering the small scale and easy diffusion of smoke targets in smoking scenarios, the spatial distance constraint weight can be appropriately increased to capture the weak spatial association between smoke and the human body; considering the drastic dynamic changes of flame targets in flame scenarios, the structural prior constraint weight can be appropriately increased to strengthen the binding relationship between flames and the human handheld area; considering the diverse types and large differences in shape of hazardous materials in hazardous material scenarios, an adaptive penalty value adjustment mechanism can be used to handle the spatial attachment patterns of different items; considering the high accuracy requirements in tool monitoring scenarios, the cost threshold can be appropriately tightened and the number of temporal validation frames increased to reduce the false alarm rate. The optimal parameter combination should be determined through cross-validation while ensuring a balance between matching recall and precision.

[0089] like Figure 2 As shown, to facilitate understanding of the working principle of the above-described target matching method for multi-target scenarios, this application embodiment also provides a specific application example of the method in a certain application scenario. In this application scenario, the above-described target matching method for multi-target scenarios mainly includes: Step 1: Obtain the target detection results from the surveillance images; The target detection results include: (1) Human target set: Each human target is output by the target detection model and includes at least: coordinates of multiple human bounding boxes: ,in , , , These represent the horizontal and vertical coordinates of the top left and bottom right corners of the box, respectively. (2) Set of face targets: the coordinates of the corresponding face bounding boxes (face targets may not be detected): (3) Each smoke target is output by the target detection model, including: smoke bounding box coordinates: .

[0090] Step 2: Calculate the center positions of the human body, face, and smoke; The center of the human body box is: ; The center of the face box is: ; The center of the smoke box is: .

[0091] Step 3: Construct the cost matrix; For each combination of human and smoke targets, construct the cost function:

[0092] in, For the first Personal goals and the first A comprehensive matching cost function between several targets to be associated (smoke targets in this application scenario); This indicates the calculation of Euclidean distance; First distance constraint An adjustable weight parameter is used to control the relative contribution of the Euclidean distance between the human body center point and the target center point to be associated in the total cost. This parameter adjusts the influence of the geometric proximity metric on the matching decision to achieve the weight allocation of spatial distance constraints in cost evaluation; the first distance constraint For the first Personal body target center point With the The center point of the smoke target The Euclidean distance function between the two targets is used to quantify the degree of spatial separation between them in the image coordinate system. The smaller the distance value, the closer the geometric positions of the candidate matching pair are, and the higher the matching rationality. The first penalty item The adjustable weight parameters are used to adjust the relative importance of the first penalty term in the total cost. By balancing the interaction between geometric distance and structural semantic constraints through weight configuration, the cost evaluation can dynamically adjust the penalty intensity for violations of structural rules according to different scenario requirements. The first penalty term is used when the center point of the target to be associated is inside the structural association area of ​​the human target, and the second preset value (usually zero or a small value) is used when it is outside. The discretization penalty mechanism encodes whether the target to be associated conforms to the semantic rules of human spatial interaction. For the second distance constraint The adjustable weight parameter controls the relative contribution of the Euclidean distance between the face center point and the center point of the target to be associated. This parameter is activated when a face target is detected, adjusting the weight ratio of the local spatial constraints of the face region in the total cost; the second distance constraint Center point of the face With the center point of the target to be associated The Euclidean distance function between the smoke target and the face region is used to measure the proximity of the smoke target and the face region. When the interaction behavior is closely related to the face region, this distance term provides a fine-grained spatial discrimination criterion.

[0093] The second penalty item The adjustable weight parameters are used to adjust the relative contribution of the second penalty term in the total cost, and the hierarchical fusion of human-level constraints and face-level constraints is achieved through weight configuration. The second penalty is to take a third preset value when the center point of the target to be associated is inside the structural association area of ​​the face target, and a fourth preset value when it is outside. This differentiated penalty strengthens the spatial constraint of the face region on the target to be associated.

[0094] The aforementioned cost function combines spatial proximity with prior knowledge of human anatomy, which increases spatial constraint information compared to simple distance matching.

[0095] Step 4: Complete the cost matrix into a square matrix; When the number of human targets and the number of smoke targets are inconsistent, the cost matrix is ​​expanded into a square matrix by introducing high-cost padding values:

[0096] in, The expanded square matrix cost matrix is ​​transformed into a square matrix structure that meets the dimensional requirements of the Hungarian algorithm by introducing virtual nodes. The original cost matrix has the following dimensions: ,in, This indicates the number of human targets detected. Indicates the number of detected targets to be associated; matrix elements Characterizing the first Personal goals and the first The matching cost between the targets to be associated; BIG is a preset constant value used to fill the matrix element region corresponding to the virtual node. Its magnitude must be much larger than the reasonable matching cost range to ensure that the cost of forming a matching pair between the real target and the virtual node is not competitive in the optimization process of minimizing the total cost, thereby preventing the algorithm from selecting virtual matches that have no physical meaning.

[0097] Step 5: Use the Hungarian algorithm to find the best matching; The Hungarian algorithm is an optimization algorithm for solving the minimum weight matching problem in bipartite graphs. The target matching method for multi-target scenarios provided in this application uses the Hungarian algorithm to solve the global optimal association between human and smoke targets. In complex multi-target scenarios, it effectively avoids the local optimum problem caused by greedy matching strategies. The core ideas of the Hungarian algorithm include: (1) reducing the cost matrix by rows and columns; (2) constructing a zero-element coverage path; (3) finding the optimal match by continuously adjusting the latent function; and (4) finally obtaining the matching scheme with the global minimum total cost. Specifically: In the preceding steps, the human body-smoke cost matrix has been constructed: The goal is to find a set of matching relationships: This ensures that each human body matches at most one smoke cloud, each smoke cloud matches at most one human body, and the total matching cost is minimized. Assume the final cost matrix used is... Its dimensions are (like (This is achieved by expanding the matrix into a square matrix by adding virtual nodes). The specific steps of the Hungarian algorithm are as follows: (1) Perform row reduction operation on the cost matrix: for each row Calculate the minimum element value in this row. And subtract the minimum value from all elements in that row, i.e., perform... ← This operation ensures that each row of the matrix produces at least one zero element and does not change the relative cost difference between any matching combinations, laying the foundation for subsequent optimization.

[0098] (2) Perform column reduction operation on the cost matrix: for each column Calculate the minimum element value in this column. And subtract the minimum value from all elements in that column, i.e., perform... ← This operation ensures that each column of the matrix contains at least one zero element. After row and column reduction, a set of zero-cost candidate edges is formed in the matrix, and all zero element positions correspond to potential optimal matching edges.

[0099] (3) Construct a zero-element bipartite graph model: Define the left node set as corresponding to multiple human targets, the right node set as corresponding to multiple targets to be associated, and the edge set as satisfying With the positions of the matrix elements determined, the positions of the zero elements are transformed into connecting edges in a bipartite graph. The problem then becomes a decisional problem: does there exist a set of non-conflicting zero-element matchings that covers all nodes?

[0100] (4) Perform minimum zero-coverage detection: find the minimum number of rows and columns required to cover all zero elements. If the minimum coverage number is equal to the matrix order, then the minimum coverage number is determined. If the minimum coverage number is less than 1, it indicates that a perfect match exists in the current set of zero elements, and we can proceed directly to step (6); If the current set of zero elements cannot form a perfect match, the matrix needs to be adjusted to generate new zero elements.

[0101] (5) Perform latent function adjustment to correct the matrix: Let the set of uncovered rows be The set of columns that are not covered is Calculate the minimum value among the uncovered elements: The adjustment rule is: subtract all uncovered elements from... Elements that overlap at row and column intersections plus The remaining covered but non-intersecting elements remain unchanged. This operation generates at least one new zero element while preserving the existing zero elements, creating conditions for the next iteration. Its expression is:

[0102] in, The new generation value after adjusting the matrix elements is the first... Line 1 The numerical result of the column elements after the latent function transformation.

[0103] (6) Extract the maximum matching of zero elements as the final solution: Solve the maximum cardinality matching problem in the current zero-cost matrix. That is, to find the matching set that covers the most nodes and does not share nodes among all zero elements, where, For binary decision variables, when matrix elements The value is 1 when selected for matching, and 0 otherwise, indicating the first... Personal goals and the first Whether a matching relationship is established between the targets to be associated. The above matching set is the globally optimal association result between human targets and targets to be associated. Each human target matches at most one target to be associated, and each target to be associated matches at most one human target, and the total matching cost reaches the global minimum.

[0104] Step Six: Matching Result Filtering and Smoking Behavior Determination; After obtaining the optimal matching result, based on the preset cost threshold Filter the matching pairs and keep only those that meet the criteria. The system filters out unreasonable matches with excessively high costs based on the matching relationships. Each valid match that passes threshold verification, corresponding to a human target and a smoke target, is determined as a smoking behavior, forming a behavior recognition result. For smoke detection results that fail to match, the system implements a hierarchical processing mechanism: when the minimum matching cost of the smoke target is higher than the first threshold, it is determined as a spatially unreasonable high-cost target and directly filtered out; when the current cost is between the first and second thresholds, the smoke target is marked as a suspected candidate and a temporal consistency verification process is triggered. Temporal consistency verification requires suspected candidates to continuously meet the matching conditions within a multi-frame image sequence of a preset number of frames. By accumulating evidence between consecutive frames, the risk of misjudgment caused by instantaneous interference in a single frame is reduced. If the suspected candidate remains stable during temporal verification, it is finally confirmed as a smoking behavior; if the verification fails or times out, it is discarded. This mechanism effectively suppresses false detections caused by environmental interference and avoids missed detections caused by single-frame matching failures.

[0105] A schematic diagram of a detection result output in the above application scenario is shown below. Figure 3 As shown in the figure, there are three human targets (person box1, person box2, and person box3) and two smoke targets (smoke box1 and smoke box2). Simultaneously, face targets (face box1 and face box2) corresponding to person box1 and person box2 are detected. The target matching method for multi-target scenarios provided in this application constructs a cost matrix that integrates the distance constraints between the human center and the smoke center, the distance constraints between the face center and the smoke center, and the penalty terms for areas inside and outside the human / face structure association region. It uses the Hungarian algorithm to globally optimize the matching of the three people and two smoke targets, ensuring a one-to-one optimal matching relationship between each smoke target and the actual smoker, avoiding multiple matchings or mismatches due to local rules. For person box3 that fails to match, a preset cost threshold can be used to correctly filter it, preventing false alarms and improving the accuracy and robustness of smoking behavior recognition in complex multi-target scenarios. When displaying the detection results, the confidence level of each detection box can also be displayed. Figure 3 Taking the shown detection boxes as an example, the detection confidence scores are 0.96 for person box1, 0.97 for person box2, 0.86 for person box3, 0.54 for smokebox1, and 0.74 for smokebox2. Furthermore, in some application scenarios, after obtaining authorization, the system can identify the person's identity based on their facial information in a face database. If the corresponding person's identity information is not identified, the system will then detect the person within the face detection box (i.e.,...). Figure 2The red detection boxes (face box1 and face box2) shown in the image are marked with: name:NoName:-1.00 (meaning that the identity information of the corresponding person was not recognized in the face database).

[0106] like Figure 4 As shown, based on the same inventive concept, this application also provides a target matching system 200 for multi-target scenarios, including a host computer 210 and an image acquisition device 220 communicatively connected to the host computer, wherein: Image acquisition device 220 is used to acquire monitoring images and send the monitoring images to host computer 210; The host computer 210 is used to acquire detection information of multiple human targets and multiple targets to be associated in the monitoring image; based on the detection information, a cost matrix is ​​constructed; the cost matrix is ​​used to characterize the matching cost between each human target and each target to be associated; the cost matrix integrates spatial distance constraints and human structure prior constraints, and the human structure prior constraints are used to penalize candidate matches where the target to be associated is located outside the structural association region of the human target; the cost matrix is ​​solved to obtain the optimal matching result between the human target and the target to be associated; based on the optimal matching result and the preset cost threshold, unreasonable matches are filtered to obtain the final matching result.

[0107] It is understood that the target matching system 200 for multi-target scenarios provided in this application embodiment can be used to execute the target matching method for multi-target scenarios provided in this application embodiment. Its implementation principle and the resulting technical effects have been described in the foregoing method embodiments. For the sake of brevity, any part not mentioned in the system embodiment can be referred to the corresponding content in any of the foregoing method embodiments.

[0108] Figure 5 This is a schematic diagram of an electronic device provided in an embodiment of this application. (Refer to...) Figure 5 The electronic device 300 includes a processor 310, a memory 320, and a communication interface 330. These components are interconnected and communicate with each other via a communication bus 340 and / or other forms of connection mechanism (not shown).

[0109] The memory 320 includes one or more (only one is shown in the figure), which may be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc. The processor 310 and other possible components may access the memory 320 to read and / or write data therein.

[0110] Processor 310 includes one or more (only one is shown in the figure), which can be an integrated circuit chip with signal processing capabilities. The processor 310 described above can be a general-purpose processor, including a central processing unit (CPU), a microcontroller unit (MCU), a network processor (NP), or other conventional processors; it can also be a special-purpose processor, including a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0111] Communication interface 330 includes one or more (only one is shown in the figure) that can be used to communicate directly or indirectly with other devices to exchange data. For example, communication interface 330 can be an Ethernet interface; it can be a mobile communication network interface, such as an interface for 3G, 4G, or 5G networks; or it can be other types of interfaces with data transmission and reception capabilities.

[0112] One or more computer program instructions may be stored in the memory 320. The processor 310 may read and run these computer program instructions to implement the target matching method for multi-target scenarios and other desired functions provided in the embodiments of this application.

[0113] Understandable. Figure 5The structure shown is for illustrative purposes only; the electronic device 300 may also include components that are more advanced than those shown. Figure 5 The more or fewer components shown, or having the same Figure 5 The different configurations shown. Figure 5 The components shown can be implemented using hardware, software, or a combination thereof. For example, electronic device 300 can be a single server (or other device with computing power), a combination of multiple servers, a cluster of a large number of servers, etc., and can be either a physical device or a virtual device.

[0114] This application also provides a computer-readable storage medium storing computer program instructions. These instructions are read and executed by a processor to perform the target matching method for multi-target scenarios provided in this application. For example, the computer-readable storage medium can be implemented as follows: Figure 5 The memory 320 in the electronic device 300, or a separate storage product (such as a USB flash drive, portable hard drive, etc.).

[0115] This application also provides a computer program product, which includes computer program instructions. These computer program instructions are read and executed by a processor to perform the target matching method for multi-target scenarios provided in this application. For example, these computer program instructions can be stored in... Figure 5 The memory 320 in the electronic device 300 is located inside the memory, or it is stored in a separate storage product (such as a USB flash drive, portable hard drive, etc.).

[0116] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0117] Furthermore, the units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0118] Furthermore, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0119] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the drawings are intended to cover non-exclusive inclusion.

[0120] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly defined.

[0121] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0122] In the description of the embodiments in this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.

[0123] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A target matching method for multi-target scenarios, characterized in that, The method includes: Acquire detection information of multiple human targets and multiple targets to be associated in surveillance images; Based on the detection information, a cost matrix is ​​constructed; wherein, the cost matrix is ​​used to characterize the matching cost between each human target and each target to be associated; the cost matrix integrates spatial distance constraints and human structure prior constraints, and the human structure prior constraints are used to penalize candidate matches where the target to be associated is located outside the structural association region of the human target; Solve the cost matrix to obtain the optimal matching result between the human target and the target to be associated; Based on the optimal matching result and the preset cost threshold, unreasonable matches are filtered out to obtain the final matching result.

2. The target matching method for multi-target scenarios according to claim 1, characterized in that, The construction of the cost matrix includes: For a human target-target pair to be associated, a spatial distance constraint component and a priori human structure constraint component are determined in the cost matrix. The spatial distance constraint component includes a first distance constraint, which quantifies the distance between the center point of the human target and the center point of the target to be associated. The priori human structure constraint component includes a first penalty term. When the center point of the target to be associated is located inside the structural association region of the human target, the first penalty term is a first preset value; when the center point of the target to be associated is located outside the structural association region, the first penalty term is a second preset value; the first preset value is less than the second preset value. The cost matrix is ​​obtained by fusing the spatial distance constraint component and the prior human structure constraint component.

3. The target matching method for multi-target scenarios according to claim 2, characterized in that, When a facial target associated with the human target is detected, the determination of the spatial distance constraint component and the prior constraint component of the human structure in the cost matrix includes: A second distance constraint is determined for the spatial distance constraint component; wherein the second distance constraint is used to quantify the distance between the center point of the target to be associated and the center point of the face; A second penalty term is determined for the prior constraint components of the human body structure; wherein, when the center point of the target to be associated is located inside the structural association region of the face target, the second penalty term is a third preset value; when the center point of the target to be associated is located outside the structural association region of the face target, the second penalty term is a fourth preset value; the third preset value is less than the fourth preset value.

4. The target matching method for multi-target scenarios according to claim 2, characterized in that, The process of fusing the spatial distance constraint component and the prior human structure constraint component to obtain the cost matrix includes: Based on the fusion weights, the spatial distance constraint component and the prior human structure constraint component are weighted and fused to obtain the cost matrix; wherein, the fusion weights are used to quantify the relative contributions of the spatial distance constraint component and the prior human structure constraint component in the cost matrix.

5. The target matching method for multi-target scenarios according to claim 1, characterized in that, Solving the cost matrix to obtain the optimal matching result between the human target and the target to be associated includes: The cost matrix is ​​reduced by columns and rows to produce at least one zero element in each row and column, thus obtaining a bipartite graph model; wherein, the bipartite graph model includes a first set of nodes representing multiple human targets and a second set of nodes representing multiple targets to be associated; By using zero-element coverage detection and matrix element iterative adjustment, the minimum cost matching scheme that covers all nodes in the bipartite graph model is determined. Extract the set of the maximum matching with zero elements corresponding to the minimum cost matching scheme to obtain the optimal matching result with the minimum global total cost.

6. The target matching method for multi-target scenarios according to any one of claims 1 to 5, characterized in that, Before solving the cost matrix, the method further includes: When the number of human targets is inconsistent with the number of targets to be associated, the cost matrix is ​​expanded into a square matrix using virtual nodes; wherein, the matrix elements corresponding to the virtual nodes are filled with preset constant values; the preset constant values ​​are greater than preset reasonable matching costs.

7. The target matching method for multi-target scenarios according to any one of claims 1 to 5, characterized in that, The step of filtering unreasonable matches based on the optimal matching result and a preset cost threshold includes: If the target to be associated fails to match, and the minimum matching cost that the target to be associated can obtain is higher than the first threshold, the target to be associated will be judged as a false detection and filtered. If the minimum matching cost corresponding to the target to be associated is between the second threshold and the first threshold, the target to be associated is marked as a suspected candidate; wherein the second threshold is less than the first threshold; The temporal consistency of the suspected candidate is verified based on a multi-frame image sequence. If the suspected candidate continuously meets the matching conditions within a preset number of frames, it is confirmed that the target to be associated has target interaction behavior with the corresponding human target.

8. The target matching method for multi-target scenarios according to any one of claims 1 to 5, characterized in that, Obtaining the final matching result includes: The optimal matching result is verified for temporal consistency in a series of consecutive frames of images. If the matching relationship in the optimal matching result continues to exist in the series of consecutive frames of images, it is confirmed that the human target corresponding to the matching relationship has target interaction behavior with the target to be associated.

9. The target matching method for multi-target scenarios according to any one of claims 1 to 5, characterized in that, The target to be associated is a smoke target.

10. A target matching system for multi-target scenarios, characterized in that, It includes a host computer and an image acquisition device that is communicatively connected to the host computer, wherein: The image acquisition device is used to acquire monitoring images and send the monitoring images to the host computer; The host computer is used to acquire detection information of multiple human targets and multiple targets to be associated in the monitoring image; construct a cost matrix based on the detection information; wherein the cost matrix is ​​used to characterize the matching cost between each human target and each target to be associated; the cost matrix integrates spatial distance constraints and prior human structure constraints, the prior human structure constraints are used to penalize candidate matches where the target to be associated is located outside the structural association region of the human target; solve the cost matrix to obtain the optimal matching result between the human target and the target to be associated; based on the optimal matching result and a preset cost threshold, filter unreasonable matches to obtain the final matching result.

11. An electronic device, characterized in that, include: A processor, a memory, and a communication bus, wherein the processor and the memory communicate with each other via the communication bus; The memory stores program instructions that can be executed by the processor, and the processor can execute the method as described in any one of claims 1 to 9 by calling the program instructions.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a computer, cause the computer to perform the method as described in any one of claims 1 to 9.

13. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method as described in any one of claims 1 to 9.