Personnel identification method and device and storage medium

By calculating multimodal sensing information and occlusion behavior index, abnormal behavior in the target monitoring site is identified, which solves the problem of low accuracy in target personnel identification in existing technologies and achieves efficient monitoring of abnormal behavior.

CN120853115BActive Publication Date: 2026-02-10E SURFING VISION TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511359810.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2026-02-10
Estimated Expiration
2045-09-23

AI Technical Summary

Technical Problem

In existing technologies, the accuracy of identifying target personnel exhibiting abnormal behavior in target monitoring locations is low, making it difficult to achieve continuous tracking and precise positioning, resulting in a high false alarm rate.

Method used

By acquiring multimodal sensor information of the target monitoring site, calculating multimodal personnel feature recognition information and site feature recognition results, and combining the occlusion behavior index, the target personnel with abnormal behavior are identified, and a multidimensional correlation model of 'people-object-site' is established to break down the correlation barrier between individual abnormal physical signs and site gathering risks.

Benefits of technology

It eliminates the need for manual patrols or single biometric identification, improving the accuracy of abnormal behavior identification, reducing false alarm rates, and enabling precise location and continuous tracking of target personnel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853115B_ABST
    Figure CN120853115B_ABST
Patent Text Reader

Abstract

The application relates to a personnel identification method and device and a storage medium, wherein the personnel identification method comprises the following steps: acquiring multi-modal sensing information under a target monitoring place; the multi-modal sensing information comprises a snapshot image sequence; according to the multi-modal sensing information, multi-modal personnel feature identification information corresponding to a snapshot person in the target monitoring place is calculated; a target shielding object in the target monitoring place is identified according to the snapshot image sequence; the shielding time of the target shielding object is detected, and a shielding behavior index is calculated based on the shielding time and a place type parameter of the target monitoring place acquired; according to the shielding behavior index, a place feature identification result is calculated; and target personnel with abnormal behaviors are identified from the snapshot personnel based on the multi-modal personnel feature identification information and the place feature identification result. Through the application, the problem that the accuracy of identification of target personnel with abnormal behaviors is low is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision, and in particular to methods, devices and storage media for people identification. Background Technology

[0002] With the continuous upgrading of the demand for intelligent monitoring, the precise identification technology for target monitoring sites has become a key research focus in the industry. However, among related technologies, traditional monitoring methods mostly rely on manual patrols or single biometric identification, and the anomaly identification is scattered and difficult to locate the core monitoring area. Therefore, it is impossible to achieve continuous tracking and precise location of target personnel, resulting in low accuracy in identifying target personnel with abnormal behavior.

[0003] Currently, no effective solution has been proposed to address the issue of low accuracy in identifying target individuals exhibiting abnormal behavior in related technologies. Summary of the Invention

[0004] This application provides a method, apparatus, and storage medium for personnel identification, which at least addresses the problem of low accuracy in identifying target personnel exhibiting abnormal behavior in related technologies.

[0005] In a first aspect, embodiments of this application provide a method for identifying persons, the method comprising:

[0006] Acquire multimodal sensing information at the target monitoring location; the multimodal sensing information includes a sequence of captured images;

[0007] Based on the multimodal sensing information, calculate the multimodal personnel feature recognition information corresponding to the captured personnel at the target monitoring location;

[0008] The target obstruction in the target monitoring location is identified based on the captured image sequence; the obstruction time of the target obstruction is detected, and the obstruction behavior index is calculated based on the obstruction time and the location type parameter of the target monitoring location;

[0009] The location feature identification result is calculated based on the occlusion behavior index;

[0010] Based on the multimodal personnel feature recognition information and the location feature recognition results, target personnel with abnormal behavior are identified from the captured personnel.

[0011] In some embodiments, calculating the occlusion behavior index based on the occlusion time and the obtained location type parameter of the target monitoring location includes:

[0012] Obtain the occlusion type of the target occlusion and determine the occlusion type weight;

[0013] Obtain a preset time decay coefficient; calculate the occlusion behavior index based on the occlusion time, the weight of the occlusion type, the location type parameter, and the time decay coefficient.

[0014] In some embodiments, determining the occlusion type weight includes:

[0015] Obtain the preset classification weights corresponding to different preset occlusion types; construct an objective function based on the preset classification weights and preset task indicators;

[0016] Collect training data at the target monitoring location;

[0017] The training data is input into the objective function for weight optimization, and the preset classification weights are adjusted to obtain the optimized classification weights for each preset occlusion type.

[0018] Based on the type of occlusion of the target occluder, the corresponding optimized classification weight is determined as the occlusion type weight.

[0019] In some embodiments, the captured image sequence includes a set of captured images from different time periods; the step of calculating the location feature recognition result based on the occlusion behavior index includes:

[0020] Calculate the number of image frames in each of the captured image sets that exhibit camouflage behavior;

[0021] Based on the number of image frames and the number of images in the captured image set, the camouflage behavior coefficient is calculated for each time period.

[0022] The location feature recognition result is calculated based on the occlusion behavior index and the camouflage behavior coefficient within each time period.

[0023] In some embodiments, after identifying the target person exhibiting abnormal behavior from the captured personnel, the method further includes:

[0024] Upon identifying the target person, obtain the target person's facial profile information;

[0025] Based on the facial profile information, the frequency of the target person's appearance in a preset area within a preset time period and the dwell time are counted; a spatiotemporal heat map is generated based on the frequency of appearance and the dwell time.

[0026] The trajectory information of the target personnel is statistically analyzed, and a group implicit association network is generated based on the trajectory information and the trajectory information of known abnormal personnel.

[0027] Based on the spatiotemporal heatmap and the group implicit association network, a target anomaly discrimination result is generated for the target personnel.

[0028] In some embodiments, after generating the target anomaly discrimination result for the target person, the method further includes:

[0029] If the target anomaly identification result indicates that the target person is identified as an abnormal person, the target anomaly keyframe is determined from the captured image sequence based on the target anomaly identification result;

[0030] Obtain a preset number of associated frames; based on the number of associated frames, extract a set of abnormal associated frames containing the target abnormal keyframe from the captured image sequence;

[0031] An anomaly evidence package is generated based on the set of anomaly-related frames.

[0032] In some embodiments, after generating the anomalous evidence package, the method further includes:

[0033] Based on the abnormal evidence package, abnormal parameters of the target personnel are detected, and historical abnormal records of the target monitoring location are obtained;

[0034] Based on the abnormal parameters and the historical abnormal records, a graded early warning score is generated, and a handling strategy result is generated based on the graded early warning score.

[0035] A standardized report is generated based on the tiered early warning score and the results of the response strategy.

[0036] In some embodiments, the multimodal person feature recognition information includes micro-expression feature information, gait feature information, and body shape feature information; the step of identifying target persons exhibiting abnormal behavior from the captured persons based on the multimodal person feature recognition information and the location feature recognition result includes:

[0037] Based on the micro-expression feature information, the gait feature information, and the body shape feature information, the abnormal behavior recognition result is calculated;

[0038] Based on the location type of the target monitoring location, weight values ​​are assigned to the abnormal behavior identification results and the location feature identification results, and the abnormal behavior identification results and the location feature identification results are weighted and fused based on the weight values ​​to obtain the target identification result; the target identification result is used to indicate the target personnel with abnormal behavior identified from the captured personnel.

[0039] Secondly, embodiments of this application provide a personnel identification device, including:

[0040] The acquisition module is used to acquire multimodal sensing information of the target monitoring location; the multimodal sensing information includes a sequence of captured images.

[0041] The personnel feature recognition module is used to calculate multimodal personnel feature recognition information corresponding to the captured personnel in the target monitoring location based on the multimodal sensing information.

[0042] The index calculation module is used to identify target obstructions in the target monitoring location based on the captured image sequence; detect the obstruction time of the target obstruction; and calculate the obstruction behavior index based on the obstruction time and the location type parameter of the target monitoring location.

[0043] The scene feature recognition module is used to calculate the scene feature recognition result based on the occlusion behavior index;

[0044] The target recognition module is used to identify target personnel with abnormal behavior from the captured personnel based on the multimodal personnel feature recognition information and the scene masking recognition results.

[0045] Thirdly, embodiments of this application provide a storage medium storing a computer program thereon, which, when executed by a processor, implements the personnel identification method as described in the first aspect above.

[0046] Compared to related technologies, the personnel identification method, apparatus, and storage medium provided in this application acquire multimodal sensing information from a target monitoring location; the multimodal sensing information includes a sequence of captured images; based on the multimodal sensing information, multimodal personnel feature identification information corresponding to the captured personnel in the target monitoring location is calculated; target obstructions in the target monitoring location are identified based on the captured image sequence; the obstruction time of the target obstructions is detected, and an obstruction behavior index is calculated based on the obstruction time and the location type parameter of the acquired target monitoring location; the location feature identification result is calculated based on the obstruction behavior index; and target personnel with abnormal behavior are identified from the captured personnel based on the multimodal personnel feature identification information and the location feature identification result.

[0047] The above methods eliminate the need for manual patrols or single biometric identification. Furthermore, a multi-dimensional correlation model of "people-objects-places" is established, breaking down the barriers between "abnormal individual physical signs" and "place clustering risks." This avoids the phenomenon of scattered anomalies making it difficult to locate core monitoring areas, thereby effectively reducing the false alarm rate and solving the problem of low accuracy in identifying target personnel with abnormal behavior.

[0048] Details of one or more embodiments of this application are set forth in the following drawings and description to make other features, objects and advantages of this application more readily apparent. Attached Figure Description

[0049] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0050] Figure 1 This is a hardware structure block diagram of a terminal for a personnel identification method according to an embodiment of this application;

[0051] Figure 2 This is a flowchart of a personnel identification method according to an embodiment of this application;

[0052] Figure 3 This is a flowchart of another person identification method according to an embodiment of this application;

[0053] Figure 4 This is a structural block diagram of a personnel identification device according to an embodiment of this application. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of this application clearer, the application is described and illustrated below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the application. All other embodiments obtained by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application. Furthermore, it is understood that although the efforts made in such a development process may be complex and lengthy, for those skilled in the art related to the content disclosed in this application, modifications to design, manufacturing, or production based on the technical content disclosed in this application are merely conventional technical means and should not be construed as insufficient disclosure of the content of this application.

[0055] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that is mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.

[0056] Unless otherwise defined, the technical or scientific terms used in this application shall have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms “a,” “an,” “an,” “the,” and similar words used in this application do not indicate quantity limitation and may indicate singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units not listed, or may include other steps or units inherent to these processes, methods, products, or devices. The terms “connected,” “linked,” “coupled,” and similar words used in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. “Multiple” used in this application means two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can represent: A alone, A and B simultaneously, and B alone. The terms “first,” “second,” “third,” etc., used in this application are merely to distinguish similar objects and do not represent a specific ordering of the objects.

[0057] The method embodiments provided in this example can be executed on a terminal, computer, or similar computing device. Taking running on a terminal as an example, Figure 1 This is a hardware structure block diagram of a terminal for a personnel identification method according to an embodiment of this application. Figure 1 As shown, a terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. Optionally, the terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the terminal described above. For example, the terminal may also include components that are larger than... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0058] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the personnel identification method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thereby implementing the aforementioned method. The memory 104 may include high-speed random access memory and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0059] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the terminal's communication provider. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0060] This embodiment provides a method for personnel identification. Figure 2 This is a flowchart of a personnel identification method according to an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps:

[0061] Step S210: Obtain multimodal sensing information at the target monitoring location; the multimodal sensing information includes a sequence of captured images.

[0062] The aforementioned multimodal sensor information is collected by multimodal sensors deployed in target monitoring locations (such as public areas, key management areas, etc.). These multimodal sensors may include on-site 3D structured light camera arrays, integrated thermal imaging modules, smart trash cans with built-in weight / metal sensors, and location access control systems.

[0063] Accordingly, the acquired multimodal sensor information is centered on the captured image sequence (composed of image frames captured continuously and at regular intervals by a high-definition camera, such as 5 frames per second, recording the complete visual trajectory of a person from entering to leaving the venue). In addition, it may include other sensor data, such as audio and skin temperature gradient analysis.

[0064] Step S220: Based on the multimodal sensing information, calculate the multimodal personnel feature recognition information corresponding to the captured personnel at the target monitoring location.

[0065] Among them, by integrating heterogeneous data collected by multimodal sensors in the target monitoring site, a three-dimensional feature profile of the person being captured is constructed.

[0066] Specifically, face detection and feature point localization can be performed first based on visible light image sequences to extract facial biometric features (such as interpupillary distance, nose width, and facial contours), while infrared thermal imaging data is combined to supplement skin temperature distribution features. Second, human pose estimation algorithms are used to analyze limb movements in captured images to extract dynamic behavioral features such as gait cycles and joint angle changes. Three-dimensional point cloud data of personnel acquired by microwave radar or lidar is used to calculate body shape parameters (such as height, shoulder width, and body proportions). Furthermore, feature vectors from different modalities are spatiotemporally aligned and correlated, and a weighted fusion model is used to eliminate the limitations of a single modality (such as face recognition failure due to lighting changes or gait analysis bias caused by occlusion). Finally, a composite feature set containing multi-dimensional information such as biometrics, behavioral patterns, and body shape features is generated. This feature set supports accurate matching of individual identities and provides a structured data foundation for subsequent abnormal behavior analysis.

[0067] The above method covers the multi-dimensional characteristics of people. Even if there is partial occlusion, it can achieve accurate characterization through the features of the unoccluded area. This effectively solves the problem that related technologies can only identify obvious abnormalities (such as fainting) and cannot capture early signs such as micro-expressions (such as pupil changes) and physiological signals (such as the amount of sweat).

[0068] Step S230: Identify target obstructions in the target monitoring area based on the captured image sequence; monitor the obstruction time of the target obstructions; and calculate the obstruction behavior index based on the obstruction time and the location type parameter of the identified target monitoring area.

[0069] In this step, by analyzing a series of continuously captured images within the target monitoring area, we accurately identify situations where individuals use obstructions (including existing facilities or temporary obstacles) to occlude their behavior, and quantify the degree of abnormality in these obstructive behaviors. This provides crucial contextualized quantitative evidence for subsequent overall area anomaly assessment and individual anomaly identification. The process is described in detail below:

[0070] First, target occlusions are identified. This involves using a 3D spatial model of the target monitoring site (such as the location and size information of fixed facilities in pre-recorded site CAD drawings) and a sequence of captured images (composed of continuously captured image frames from a high-definition camera, recording the complete visual trajectory of a person from entering to leaving the site). Deep learning target detection algorithms (such as YOLOv8 and Faster R-CNN) are used to analyze each frame of the images to identify objects that obscure a person's body or face – these objects are called "target occlusions." These include fixed obstacles such as pillars, trash cans, and large green plants, as well as temporary obstacles such as cardboard boxes, shopping carts, and umbrellas. For example, in a public indoor scene, the algorithm can detect the bounding box (e.g., coordinates (x1, y1, x2, y2)) of a "table" from consecutive image frames. By matching it with the 3D model of the location, if it detects a person using the "table" to obscure their face, it can be labeled as "fixed obscuring object - table". If it detects a person using an "umbrella" to obscure their face, it can be labeled as "temporary obscuring object - umbrella". This process must ensure the accuracy of obscuring object recognition (avoiding misidentifying a person's own clothing as an obscuring object) and real-time performance (processing time per frame not exceeding 0.1 seconds to meet real-time monitoring requirements).

[0071] Next, the duration of each occlusion event is monitored. Specifically, for each captured person, their position is tracked in the image sequence using a multi-target tracking algorithm (such as DeepSORT). When the overlap rate between the bounding box of the target occluder and the bounding box of the person exceeds a preset threshold (e.g., 30%, meaning the occluder covers 30% or more of the person's body or face), the occlusion start time is recorded (e.g., frame 15, corresponding to 1 minute and 3 seconds in actual time). When the overlap rate is below the threshold (or the person completely leaves the occluder area), the occlusion end time is recorded (e.g., frame 55, corresponding to 15 minutes and 3 seconds in actual time). The time difference between the two is the occlusion time of the target occluder on the person (e.g., 15 minutes and 3 seconds - 1 minute and 3 seconds = 14 minutes). For example, if a person is walking in a monitored area and is blocked by a pillar for 14 minutes, the algorithm will accurately record this time period and set a threshold to filter short-term interference. For example, less than 5 seconds is considered invalid, in order to ensure the continuity of the blocking time (not missing the blocking status of intermediate frames) and accuracy (excluding the misjudgment of instantaneous blocking, such as the brief overlap when people brush past each other).

[0072] It should also be noted that the above-mentioned venue type parameters are quantitative indicators predefined or trained through machine learning based on the attributes of the target monitored venue (venue function, location, personnel flow patterns, historical occlusion data), used to reflect the boundaries of "normal occlusion behavior" in this type of venue. These venue type parameters can be classified according to the venue's risk level. Specific methods for obtaining these parameters include: statistically analyzing the intentional occlusion rate of different venues based on historical data; this indicator characterizes the proportion of target personnel intentionally using obstructions under normal circumstances; for example, in residential areas, the intentional occlusion rate is 5%, meaning approximately 5 out of 100 people will intentionally use pillars or other obstructions; or statistically analyzing the maximum permissible normal occlusion time for different venues; or dynamically adjusting the personnel flow rate adjustment coefficient based on the current personnel flow rate of the venue, and combining this with the intentional occlusion rate, occlusion time threshold, personnel flow rate adjustment coefficient, etc., to statistically determine the venue risk level, thereby determining the above-mentioned venue type parameters.

[0073] Finally, by combining the aforementioned occlusion time and location type parameters, an occlusion behavior index is calculated to quantify the abnormality of the occlusion behavior. It is evident that quantifying the key indicator of the abnormal behavior of a target individual "deliberately using location facilities (such as pillars) to obscure their face" based on the occlusion behavior index allows for the differentiation between "accidental occlusion" (such as being obscured by a pillar while passing by) and "deliberate occlusion" (such as actively using a pillar to conceal one's identity and behavior) by combining "occlusion time" (behavior duration) and "location type parameters" (scene adaptability), thereby improving the accuracy of target individual identification.

[0074] Step S240: Calculate the location feature recognition result based on the occlusion behavior index.

[0075] Specifically, the occlusion behavior index is correlated with the inherent characteristics of the location to output the location feature recognition results, providing contextual information for subsequent abnormal behavior recognition.

[0076] Step S250: Based on multimodal personnel feature recognition information and location feature recognition results, identify target personnel with abnormal behavior from the captured personnel.

[0077] The system dynamically weights and fuses the multimodal personnel feature recognition information of each captured person with the location feature recognition results (location anomaly score) to generate a comprehensive personnel anomaly index. Based on the comparison between this comprehensive index and a preset threshold, it determines whether there are target personnel exhibiting abnormal behavior. Such target personnel typically exhibit non-routine and irregular activity patterns, and their environment may contain abnormal personnel interactions or hidden objects.

[0078] For example, if a person's multimodal features show "abnormal pupil dilation + unsteady gait" (person feature anomaly), and the feature result of the location is "highly anomaly" (location anomaly), then the generated comprehensive anomaly index for the person, calculated through fusion, indicates an extremely high probability of "anomaly," and this person will be marked as a target person (an anomaly object requiring close attention). This step achieves synergy between "individual features" and "scene context," significantly improving the accuracy and reliability of anomaly identification.

[0079] Through steps S210 to S250, multimodal personnel feature recognition information generated from multimodal sensing information and location feature recognition results generated based on at least the occlusion behavior index corresponding to the target occlusion object in the target monitoring location are used to comprehensively analyze whether there are target personnel with abnormal behavior. Therefore, it is not necessary to rely on manual patrols or single biometric identification. Furthermore, a multidimensional correlation model of "person-object-location" is established on this basis, which breaks down the correlation barrier between "individual abnormal physical signs" and "location clustering risk", avoiding the phenomenon that scattered abnormalities make it difficult to locate the core monitoring area, thereby effectively reducing the false alarm rate and solving the problem of low accuracy in identifying target personnel with abnormal behavior.

[0080] In some embodiments, the calculation of the occlusion behavior index based on the occlusion time and the obtained location type parameter of the target monitoring location may further include the following steps:

[0081] Obtain the occlusion type of the target occlusion and determine the occlusion type weight; obtain the preset time decay coefficient; calculate the occlusion behavior index based on the occlusion time, occlusion type weight, location type parameter and time decay coefficient.

[0082] In this embodiment, the process of calculating the occlusion behavior index also incorporates a dynamic fusion and attenuation mechanism of multi-dimensional parameters. Specifically, firstly, the type of occlusion (such as static facilities, temporary obstacles, etc.) is identified through a target detection algorithm, and a preset risk weight is assigned to each type (for example, the weight of temporarily placed items is higher than that of fixed facilities); secondly, a time decay coefficient is introduced, which decreases non-linearly with the increase of occlusion time (such as using an exponential decay model) to avoid over-calculation of long-term occlusion; finally, the occlusion time, the weight of the occlusion type, the location type parameter (such as higher weight for high-risk areas), and the time decay coefficient are integrated through a weighted summation or product model to generate a quantitative index reflecting the potential risk of occlusion behavior. This index achieves an environmentally adaptive quantitative assessment of the degree of abnormality of occlusion behavior by dynamically adjusting parameters (such as real-time updates of location type weights) and using an attenuation mechanism (such as rapid reduction of impact after exceeding a threshold).

[0083] One method for calculating the aforementioned occlusion behavior index can be illustrated by the following formula:

[0084] ;

[0085] In the above formula, This index represents the occlusion behavior. n is the number of currently identified target occlusions. Used to represent the weight of the occlusion type of the i-th target occlusion; Used to represent the duration of a single occlusion action corresponding to the i-th target occluder; Used to represent the weight of the venue type (i.e., the venue type parameter mentioned above). Used to represent the time decay coefficient (to suppress overfitting of historical data).

[0086] Through the above embodiments, by introducing occlusion type weights to distinguish the potential risks of different occlusions, misjudgments caused by homogenization are avoided. Secondly, the time decay coefficient, through a non-linear decay model (such as an exponential function), ensures that the impact of recent occlusion behavior on the index is significantly higher than that of historical data, ensuring that the assessment results reflect the real-time status. Finally, by combining location type parameters (such as weighting high-risk areas), scenario-based adaptation of risk assessment is achieved. The synergistic effect of these three elements enables the occlusion behavior index to have dynamic weight adjustment capabilities, which can highlight key occlusion behaviors in high-risk scenarios and filter out occasional occlusions in low-risk areas, ultimately improving the accuracy of abnormal behavior identification. At the same time, the time decay mechanism reduces redundant data interference and optimizes the efficiency of monitoring resource allocation.

[0087] In some embodiments, the determination of the occlusion type weight may further include the following steps:

[0088] Obtain the preset classification weights corresponding to different preset occlusion types; construct an objective function based on the preset classification weights and preset task indicators; collect training data in the target monitoring location; input the training data into the objective function for weight optimization processing, and adjust the preset classification weights to obtain the optimized classification weights for each preset occlusion type; determine the corresponding optimized classification weights as occlusion type weights based on the occlusion type of the target occlusion.

[0089] Specifically, firstly, predefined classification weights for different types of occlusions are defined as the starting point for model learning. For example, the predefined classification weight for a pillar is 0.1, and the initial classification weight for an umbrella is 0.2, etc. Then, historical surveillance video frames of the target monitoring location are collected as training data, and the true type labels of the occlusions in each frame are labeled (e.g., labeling a pillar obscuring a face in a certain frame). Subsequently, combined with the core task indicators of the location, such as the false negative rate and false positive rate of target detection, or the accuracy and recall rate of target personnel identification, the predefined classification weights are correlated with the task indicators to construct an objective function. Essentially, this is a quantitative model of "weighted occlusion impact × task loss," used to evaluate the rationality of the weights.

[0090] Next, real training data (including occlusion type labels, occlusion degree, task results, etc.) from the target monitoring location is collected and used as input to the objective function. Machine learning optimization algorithms (such as gradient descent and Adam) are used to adjust the preset classification weights, gradually converging the objective function to its optimum. This minimizes the negative impact of "high-weight occlusion types" on the task metrics, yielding optimized classification weights for each occlusion type. Finally, in practical applications, when a target occlusion (such as a pedestrian occlusion in an image frame) is detected, its corresponding optimized classification weight is directly retrieved as a quantification of the occlusion's impact on the task (e.g., prioritizing pedestrian occlusion to reduce the risk of missed detections).

[0091] The above embodiments provide a data-driven weight optimization mechanism that enables scenario-based adaptive weighting of occlusion types. This process corrects empirical weights using actual scenario data, making the weights of occlusion types more consistent with the distribution characteristics of occlusions in the target location. This improves the accuracy of subsequent occlusion behavior index calculations and avoids misjudgments or omissions caused by fixed weights.

[0092] In some embodiments, the above-mentioned image capture sequence includes a set of images captured within different time periods; the above-mentioned calculation of the location feature recognition result based on the occlusion behavior index may further include the following steps:

[0093] Calculate the number of image frames exhibiting camouflage behavior in each captured image set; based on the number of image frames and the number of images in the captured image set, calculate the camouflage behavior coefficient for each time period; and calculate the location feature recognition result based on the occlusion behavior index and the camouflage behavior coefficient for each time period.

[0094] Specifically, for image sets captured at different time periods within the target monitoring location (e.g., hourly or daily image sets), the number of image frames exhibiting camouflage behavior (e.g., frames where faces or identity features are obscured by masks, sunglasses, hats, etc.) is counted. Then, this number of camouflaged frames is divided by the total number of images in the corresponding image set to obtain the camouflage behavior coefficient for each time period. This coefficient is a quantitative indicator of the frequency of camouflage behavior; a higher value indicates more prevalent camouflage behavior during that time period. These camouflage behavior coefficients are then combined with a pre-calculated occlusion behavior index. The specific combination method can be designed according to business needs, such as weighted summation, product, or more complex functions, to comprehensively consider the impact of both on the task, ultimately yielding the location feature recognition result. This result is a comprehensive assessment of the "frequency of camouflage behavior" and "occlusion situation" within the location and can be used to guide subsequent tasks.

[0095] More specifically, calculate the number of camouflage images during the day ( The total number of images captured during the day ( The percentage of camouflage images at night and the number of camouflage images at night ( The total number of photos captured at night ( The proportions in the formula are weighted and fused by multiplying each factor by the occlusion behavior index; as shown in the following formula:

[0096] ;

[0097] In the above formula, Used to represent the results of the above-mentioned location feature identification; , These are used to represent the weight values ​​corresponding to the daytime and nighttime time periods, respectively. , The parameters can be fine-tuned according to different locations.

[0098] Through the above embodiments, focusing on location feature identification, by quantifying the frequency of camouflage behavior and combining it with occlusion conditions, the accurate identification and evaluation of location features can be achieved, thereby effectively improving the accuracy of target personnel identification.

[0099] In some embodiments, after identifying target individuals exhibiting abnormal behavior from the captured personnel, the personnel identification method may further include the following steps:

[0100] Upon identifying a target person, obtain their facial profile information; based on the facial profile information, count the frequency of the target person's appearance in a preset area within a preset time period and the dwell time; generate a spatiotemporal heatmap based on the frequency of appearance and dwell time; count the target person's trajectory information, and generate a group implicit association network based on the trajectory information and the trajectory information of known abnormal persons; and generate a target anomaly discrimination result for the target person based on the spatiotemporal heatmap and the group implicit association network.

[0101] Once a target person is identified (e.g., confirmed through facial detection / recognition technology), their facial profile information (including historical activity records and identity features) is retrieved first. Based on this profile, the frequency of their appearance in a preset area (e.g., a community, public entertainment venue, or core monitoring area) and their dwell time (cumulative duration of stay in the area) within a preset time period (e.g., the last 7 days) are calculated. These two indicators together reflect the target person's "activity dependence" on the area (e.g., frequent appearances and long stays may indicate a close connection with the area). Next, these two indicators are combined with a time-space dimension to generate a spatiotemporal heat map (e.g., using time as the horizontal axis and area as the vertical axis, using color intensity to represent the activity intensity of the target person in different time periods and areas, visually displaying their activity hotspots, such as "frequently appearing around a public entertainment venue every Friday from 6 PM to 8 PM").

[0102] At the same time, the trajectory information of the target personnel (i.e., their movement paths within a preset time period, such as the route from home to a public entertainment venue) is statistically analyzed, and these trajectories are compared with the historical trajectory information of known abnormal personnel to mine the similarity of the two trajectories (such as both frequently passing through a certain alley or both appearing in a certain high-risk location at the same time period). This generates a group implicit association network (using nodes to represent personnel, edges to represent trajectory similarity, and the weight of the edges to reflect the strength of the association; for example, if the trajectory similarity between the target personnel and a known abnormal person reaches 80%, then there is a high-weight edge between the two). Through this network, implicit associations between the target personnel and abnormal groups can be discovered (indirect contact but with highly overlapping activity paths).

[0103] Finally, the spatiotemporal heatmap and the group implicit association network are combined to generate the target anomaly discrimination result: the spatiotemporal heatmap (individual activity patterns) can reveal whether the target person is frequently active in the preset anomaly sensitive area / time period (such as known anomalous personnel gathering areas, remote road sections from 2 am to 4 am) (e.g., the heatmap shows that the target person stays for more than 2 hours around an entertainment venue every Friday night, and this area is a hotspot activity area for anomalous personnel); the group implicit association network (group association) can verify whether the target person has a trajectory association with known anomalous personnel (e.g., the network shows that the trajectory similarity between the target person and 3 known anomalous personnel all exceed 70%); after combining the two, if the target person's "individual activity hotspot" and "abnormal group activity hotspot" highly overlap, and the "group association strength" reaches the preset threshold (e.g., edge weight ≥ 0.6), the anomaly discrimination result will mark the target person as a high-risk anomaly, otherwise it will be low-risk.

[0104] In short, the above embodiments, through a two-dimensional analysis combining visualization of individual spatiotemporal activity patterns (spatiotemporal heatmap) and mining of implicit group associations (association network), avoid misjudgment based on a single indicator (such as only looking at the frequency of occurrence), and achieve comprehensive and accurate identification of the abnormal status of target personnel. For example, if a target personnel appears frequently and stays for a long time in a preset area (spatiotemporal heatmap shows high activity), and their trajectory is highly similar to that of known abnormal personnel (group association network shows strong association), then their abnormal risk will be significantly increased, providing a reliable basis for subsequent handling (such as key monitoring and further investigation).

[0105] In some embodiments, after generating the target anomaly discrimination result for the target personnel, the personnel identification method may further include the following steps:

[0106] When the target anomaly identification result indicates that the target person is identified as an abnormal person, the target anomaly keyframe is determined from the captured image sequence based on the target anomaly identification result; a preset number of associated frames is obtained; based on the number of associated frames, an abnormal associated frame set containing the target anomaly keyframe is extracted from the captured image sequence; and an anomaly evidence package is generated based on the abnormal associated frame set.

[0107] When the target anomaly detection result indicates that the target person is abnormal, the system will locate the target anomaly key frame from the captured image sequence (i.e., the core image frame that best reflects the person's abnormal behavior / characteristics, such as the moment of disguise, the key moment of contact with the abnormal person, or a typical frame appearing in the abnormal area); then obtain the preset number of associated frames (a pre-set number of consecutive frames that must include the context before and after the key frame, used to completely reconstruct the abnormal behavior process, such as 50 frames before and after the key frame, for a total of 100 frames); based on this number of associated frames, extract the set of abnormal associated frames containing the target anomaly key frame from the captured image sequence (i.e., the key frame and a specified number of consecutive frames before and after it, forming a complete behavior timeline, such as from the preparatory actions before the start of the abnormal behavior to the departure process after it ends).

[0108] Finally, the abnormal frame set is organized in chronological order, and the abnormal features of the key frames are labeled (such as "Frame 120: Target person wearing sunglasses to cover face", "Frame 150: Shaking hands with known abnormal person"), and image metadata (such as shooting time, location, camera number) is added to generate an abnormal evidence package. This evidence package is a complete and traceable chain of evidence of abnormal behavior that can be used for subsequent investigations, law enforcement or judicial procedures. It contains core evidence of key moments and also covers the context of the behavior process, ensuring the authenticity and persuasiveness of the evidence.

[0109] In some embodiments, after generating the abnormal evidence package, the personnel identification method may further include the following steps:

[0110] Based on the abnormal evidence package, abnormal parameters of the target personnel are detected, and historical abnormal records of the target monitoring site are obtained; based on the abnormal parameters and historical abnormal records, a graded early warning score is generated, and a response strategy result is generated based on the graded early warning score; based on the graded early warning score and the response strategy result, a standardized report is generated.

[0111] First, abnormal parameters of the target personnel are extracted from the aforementioned abnormal evidence package, such as the frequency of camouflage behavior, the duration of stay in a specific area, and the similarity of their movement trajectory to known abnormal objects. Simultaneously, historical abnormal records of the target monitoring location are retrieved, such as similar abnormal events that occurred at the location in the past and their handling results (e.g., "In historical events with similar camouflage + trajectory abnormalities, 70% require further verification"). Next, the current abnormal parameters are compared and analyzed with historical records. For example, if the current trajectory similarity is higher than the historical high-risk threshold by 60%, or the number of camouflage instances exceeds the historical average by 1.5 times, a weighted calculation is used to generate a graded early warning score. This intelligent early warning graded mechanism can be specifically: Level 1 warning: the target personnel trigger abnormal identification 3 times within a week at the same location + historical abnormal records of the location; Level 2 warning: a single person exhibits 4 abnormal micro-expressions and a high index of location occlusion behavior.

[0112] Subsequently, the corresponding handling strategy is matched according to the scoring level. For example, high risk corresponds to "strengthening real-time monitoring and notifying relevant personnel to intervene and verify," medium risk corresponds to "continuously tracking the person for 72 hours," and low risk corresponds to "including in the regular watch list." Finally, the graded early warning scores, handling strategies, and information in the abnormal evidence package (such as keyframe annotations, related frame time sequences, and abnormal feature descriptions) are integrated to generate a standardized report (with a unified format, covering an overview of the abnormal situation, evidence details, scoring basis, and handling recommendations). This ensures that the information is complete and logically clear, facilitating subsequent archiving, storage, and processing. This effectively solves problems such as high false alarm rates, response delays, and privacy leakage risks in related technologies, and improves the efficiency of discovering core monitoring sites.

[0113] In some embodiments, the aforementioned multimodal personnel feature recognition information includes micro-expression feature information, gait feature information, and body shape feature information. Specifically, the multi-dimensional feature extraction mechanism may include the following steps: (1) Dynamic face enhancement: fusion of face + gait + body shape three-factor verification; (2) Micro-expression analysis: face: constructing a 12-dimensional feature vector (pupil contraction frequency, blink interval, etc.), body shape signal: thermal imaging detection of nose tip temperature fluctuation (temperature rise > 2℃ after use of certain controlled substances), BMI calculation (images after cropping the face are uniformly adjusted to a fixed size of 224×224, and the area and weight are calculated to be 4.48 kg / m²); (3) Behavioral map construction: abnormal action library: such as repeatedly picking at the face, scratching, cutting nails, picking pimples and other meaningless actions, accompanied by restlessness, violent shaking, neck shaking, scratching, curling up the body, unsteady gait, and uncoordinated movements. As can be seen, compared with the identification methods in related technologies that can only identify obvious abnormalities (such as fainting) and cannot capture early signs such as micro-expressions (pupil changes) and physiological signals (amount of sweating), this embodiment, by constructing a multimodal data fusion analysis system, refines the behavioral analysis, can promptly identify target personnel with abnormal behavior, and effectively reduce false alarm rate and response delay.

[0114] Based on this, the above-mentioned method of identifying target individuals exhibiting abnormal behavior from captured individuals, based on multimodal personnel feature recognition information and location feature recognition results, may further include the following steps:

[0115] Based on micro-expression features, gait features, and body shape features, the abnormal behavior recognition result is calculated. One calculation formula is shown below:

[0116] ;

[0117] ;

[0118] In the above formula, Used to represent the results of abnormal behavior identification; Used to represent micro-expression feature information, and to identify whether there are abnormalities through facial expressions; Used to represent gait feature information; Used to represent body shape characteristics; These are used to represent the weights of each feature. It should also be noted that the weights are adjusted based on the characteristics of the main controlled substances prevalent in the local area (such as a common type of psychoactive substance requiring regulation). For example, if a locally prevalent controlled substance is likely to cause users to exhibit characteristics such as "frequent trembling" or "unfocused gaze," then the weights of features such as "body trembling" and "eye state" in the model will be increased (i.e., these features will be given more importance during identification), thereby improving the accuracy of identifying users of such substances.

[0119] Next, based on the location type of the target monitoring location, weight values ​​are assigned to the abnormal behavior identification results and the location feature identification results, and the abnormal behavior identification results and the location feature identification results are weighted and fused based on the weight values ​​to obtain the target identification results; the target identification results are used to indicate the target personnel with abnormal behavior identified from the captured personnel.

[0120] The calculation process described above is illustrated in the following formula:

[0121] ;

[0122] In the above formula, Used to represent the above target identification results; , These are used to represent the weight values ​​corresponding to the abnormal behavior recognition results and the location feature recognition results, respectively. , The parameters can be fine-tuned according to different locations.

[0123] The above embodiments take into account both individual behavior and location environment information, and can ultimately accurately identify target individuals with abnormal behavior from the captured personnel.

[0124] The present application will be described and illustrated in detail below through specific embodiments. Figure 3 This is a flowchart of another person identification method according to an embodiment of this application, such as... Figure 3 As shown, the process includes the following steps:

[0125] Step S301: Collect multimodal data. This involves deploying multimodal sensors, such as a 3D structured light camera array, a thermal imaging module, smart trash cans, and access control systems, at the target monitoring location to collect multimodal data.

[0126] Step S302, Feature Extraction. That is, based on the multimodal data collected above, multi-dimensional feature extraction processing is performed, including micro-expression analysis, gait recognition, and abnormal action recognition.

[0127] Step S303: Output the anomaly detection result and determine whether it is abnormal. This includes: Location profile generation: Occlusion behavior index = Σ Occlusion type weight × Duration × Location type weight.

[0128] Step S304: If the anomaly detection result output in step S303 indicates that the data is normal, then the data is archived.

[0129] Step S305: If the anomaly detection result output in step S303 indicates an anomaly, then perform spatiotemporal correlation analysis. This includes: generating a spatiotemporal heatmap: areas with nighttime activity >70% and personnel dwell time >45 minutes are marked with a red alert; generating a group correlation network: establishing implicit associations with personnel whose trajectories overlap >80%, appear simultaneously >3 times, and have a history of anomalies.

[0130] Step S306: Based on the above analysis results, risk scoring and early warning classification are performed, and a response decision is generated. Finally, a standardized report (including risk score, related locations, and response recommendations) is generated based on this information.

[0131] It should be noted that the steps shown in the above process or in the flowchart of the accompanying figures can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0132] This embodiment also provides a personnel identification device for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the terms "module," "unit," "subunit," etc., can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0133] Figure 4 This is a structural block diagram of a personnel identification device according to an embodiment of this application, such as... Figure 4 As shown, the device includes: an acquisition module 10, a personnel feature recognition module 20, an index calculation module 30, a scene feature recognition module 40, and a target recognition module 50; wherein:

[0134] The acquisition module 10 is used to acquire multimodal sensing information of the target monitoring location, including a sequence of captured images. The personnel feature recognition module 20 is used to calculate multimodal personnel feature recognition information corresponding to the captured personnel in the target monitoring location based on the multimodal sensing information. The index calculation module 30 is used to identify target occlusions in the target monitoring location based on the captured image sequence, detect the occlusion time of the target occlusions, and calculate the occlusion behavior index based on the occlusion time and the acquired location type parameter of the target monitoring location. The scene feature recognition module 40 is used to calculate the location feature recognition result based on the occlusion behavior index. The target recognition module 50 is used to identify target personnel with abnormal behavior from the captured personnel based on the multimodal personnel feature recognition information and the scene occlusion recognition result.

[0135] In some embodiments, the personnel identification device further includes a discrimination module; this discrimination module is used to acquire the facial profile information of the target person when the target person is identified; based on the facial profile information, to count the number of times the target person appears in a preset area within a preset time period and the dwell time; to generate a spatiotemporal heat map based on the number of times the target person appears and the dwell time; to count the trajectory information of the target person, and to generate a group implicit association network based on the trajectory information and the acquired trajectory information of known abnormal persons; and to generate a target anomaly discrimination result for the target person based on the spatiotemporal heat map and the group implicit association network.

[0136] In some embodiments, the discrimination module is further configured to, when the target anomaly discrimination result indicates that the target person is judged as an abnormal person, determine the target abnormal key frame from the captured image sequence according to the target anomaly discrimination result; obtain a preset number of associated frames; extract an abnormal associated frame set containing the target abnormal key frame from the captured image sequence based on the number of associated frames; and generate an abnormal evidence package based on the abnormal associated frame set.

[0137] In some embodiments, the discrimination module is further configured to detect abnormal parameters of the target personnel based on the abnormal evidence package and obtain historical abnormal records of the target monitoring site; generate a graded early warning score based on the abnormal parameters and historical abnormal records, and generate a disposal strategy result based on the graded early warning score; and generate a standardized report based on the graded early warning score and the disposal strategy result.

[0138] It should be noted that the above modules can be functional modules or program modules, and can be implemented by software or hardware. For modules implemented by hardware, the above modules can reside in the same processor; or the above modules can be located in different processors in any combination. Specific examples in this embodiment can be found in the examples described in the above embodiments and optional implementations, and will not be repeated in this embodiment.

[0139] This embodiment also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0140] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.

[0141] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:

[0142] S1, acquire multimodal sensing information of the target monitoring site; the multimodal sensing information includes the captured image sequence.

[0143] S2, based on the multimodal sensing information, calculate the multimodal personnel feature recognition information corresponding to the captured personnel at the target monitoring location.

[0144] S3, identify target obstructions in the target monitoring area based on the captured image sequence; detect the obstruction time of the target obstruction, and calculate the obstruction behavior index based on the obstruction time and the location type parameter of the target monitoring area.

[0145] S4. Calculate the location feature recognition result based on the occlusion behavior index.

[0146] S5 identifies target individuals exhibiting abnormal behavior from captured personnel based on multimodal personnel feature recognition information and location feature recognition results.

[0147] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated here.

[0148] Furthermore, in conjunction with the personnel identification methods in the above embodiments, this application embodiment can provide a storage medium for implementation. This storage medium stores a computer program; when executed by a processor, the computer program implements any of the personnel identification methods in the above embodiments.

[0149] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0150] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0151] Those skilled in the art should understand that the technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments have been described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0152] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A method for identifying persons, characterized in that, The method includes: Acquire multimodal sensing information at the target monitoring location; the multimodal sensing information includes a sequence of captured images; Based on the multimodal sensing information, calculate the multimodal personnel feature recognition information corresponding to the captured personnel at the target monitoring location; The target obstruction in the target monitoring location is identified based on the captured image sequence; the obstruction time of the target obstruction is detected, and the obstruction behavior index is calculated based on the obstruction time and the location type parameter of the target monitoring location; The location feature identification result is calculated based on the occlusion behavior index; Based on the multimodal personnel feature recognition information and the location feature recognition results, target personnel with abnormal behavior are identified from the captured personnel; Upon identifying the target person, obtain the target person's facial profile information; Based on the facial profile information, the frequency of the target person's appearance in a preset area within a preset time period and the dwell time are counted; a spatiotemporal heat map is generated based on the frequency of appearance and the dwell time. The trajectory information of the target personnel is statistically analyzed, and a group implicit association network is generated based on the trajectory information and the trajectory information of known abnormal personnel. Based on the spatiotemporal heatmap and the group implicit association network, a target anomaly discrimination result is generated for the target personnel.

2. The personnel identification method according to claim 1, characterized in that, The calculation of the occlusion behavior index based on the occlusion time and the obtained location type parameters of the target monitoring location includes: Obtain the occlusion type of the target occlusion and determine the occlusion type weight; Obtain a preset time decay coefficient; calculate the occlusion behavior index based on the occlusion time, the weight of the occlusion type, the location type parameter, and the time decay coefficient.

3. The personnel identification method according to claim 2, characterized in that, The determination of the weight of the occlusion type includes: Obtain the preset classification weights corresponding to different preset occlusion types; construct an objective function based on the preset classification weights and preset task indicators; Collect training data at the target monitoring location; The training data is input into the objective function for weight optimization, and the preset classification weights are adjusted to obtain the optimized classification weights for each preset occlusion type. Based on the type of occlusion of the target occluder, the corresponding optimized classification weight is determined as the occlusion type weight.

4. The personnel identification method according to claim 1, characterized in that, The captured image sequence includes sets of captured images from different time periods; The step of calculating the location feature identification result based on the occlusion behavior index includes: Calculate the number of image frames in each of the captured image sets that exhibit camouflage behavior; Based on the number of image frames and the number of images in the captured image set, the camouflage behavior coefficient is calculated for each time period. The location feature recognition result is calculated based on the occlusion behavior index and the camouflage behavior coefficient within each time period.

5. The personnel identification method according to claim 1, characterized in that, After generating the target anomaly discrimination result for the target personnel, the method further includes: If the target anomaly identification result indicates that the target person is identified as an abnormal person, the target anomaly keyframe is determined from the captured image sequence based on the target anomaly identification result; Obtain a preset number of associated frames; based on the number of associated frames, extract a set of abnormal associated frames containing the target abnormal keyframe from the captured image sequence; An anomaly evidence package is generated based on the set of anomaly-related frames.

6. The personnel identification method according to claim 5, characterized in that, After generating the abnormal evidence package, the method further includes: Based on the abnormal evidence package, abnormal parameters of the target personnel are detected, and historical abnormal records of the target monitoring location are obtained; Based on the abnormal parameters and the historical abnormal records, a graded early warning score is generated, and a handling strategy result is generated based on the graded early warning score. A standardized report is generated based on the tiered early warning score and the results of the response strategy.

7. The personnel identification method according to any one of claims 1 to 6, characterized in that, The multimodal human feature recognition information includes micro-expression feature information, gait feature information, and body shape feature information; The step of identifying target individuals exhibiting abnormal behavior from the captured individuals based on the multimodal personnel feature recognition information and the location feature recognition results includes: Based on the micro-expression feature information, the gait feature information, and the body shape feature information, the abnormal behavior recognition result is calculated; Based on the location type of the target monitoring location, weight values ​​are assigned to the abnormal behavior identification results and the location feature identification results, and the abnormal behavior identification results and the location feature identification results are weighted and fused based on the weight values ​​to obtain the target identification result; the target identification result is used to indicate the target personnel with abnormal behavior identified from the captured personnel.

8. A personnel identification device, characterized in that, include: The acquisition module is used to acquire multimodal sensing information of the target monitoring location; the multimodal sensing information includes a sequence of captured images. The personnel feature recognition module is used to calculate multimodal personnel feature recognition information corresponding to the captured personnel in the target monitoring location based on the multimodal sensing information. The index calculation module is used to identify target obstructions in the target monitoring location based on the captured image sequence; detect the obstruction time of the target obstruction; and calculate the obstruction behavior index based on the obstruction time and the location type parameter of the target monitoring location. The scene feature recognition module is used to calculate the scene feature recognition result based on the occlusion behavior index; The target recognition module is used to identify target personnel with abnormal behavior from the captured personnel based on the multimodal personnel feature recognition information and the location feature recognition results; The discrimination module is used to obtain the facial profile information of the target person when the target person is identified; based on the facial profile information, it counts the number of times the target person appears in a preset area within a preset time period and the dwell time. A spatiotemporal heatmap is generated based on the occurrence frequency index and the residence time; The trajectory information of the target personnel is statistically analyzed. Based on the trajectory information and the trajectory information of known abnormal personnel, a group implicit association network is generated. Based on the spatiotemporal heat map and the group implicit association network, a target anomaly discrimination result for the target personnel is generated.

9. A storage medium, characterized in that, The storage medium stores a computer program, wherein the computer program is configured to execute the personnel identification method according to any one of claims 1 to 7 when it is run.

Citation Information

Patent Citations

  • Personnel abnormal behavior identification method based on video image

    CN117496586A