Target perception and behavior analysis method and device, computer equipment and storage medium

By integrating target detection, identity recognition, and behavior analysis technologies, the system solves the problems of multi-task collaboration and adaptability to complex environments in home care scenarios, enabling comprehensive and proactive home care and improving recognition accuracy and stability.

CN121963272APending Publication Date: 2026-05-01SHENZHEN XIAOPAI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN XIAOPAI TECHNOLOGY CO LTD
Filing Date
2026-01-04
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing home care solutions cannot meet the diverse and intelligent needs. Traditional human caregivers have limited energy, and ordinary monitoring equipment lacks the ability to actively identify and warn, making it difficult to cope with emergencies in complex home environments. Furthermore, multi-camera collaborative identification has the problem of identity confusion.

Method used

Employing target detection, identity recognition, trajectory tracking, and behavior analysis technologies, this system integrates lightweight human and face detection models with multimodal large models for behavior and emotion analysis. By using feature filtering and merging strategies, it achieves identity management and trajectory stability, providing comprehensive proactive care.

Benefits of technology

It improves the real-time performance and identification accuracy of home care solutions, solves the problems of multi-task coordination and long-term stable tracking in complex environments, and ensures comprehensive safety monitoring of family members.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963272A_ABST
    Figure CN121963272A_ABST
Patent Text Reader

Abstract

The invention discloses a target perception and behavior analysis method, which comprises the steps of detecting a human body and a human face in a picture, and extracting a human body and human face feature set; performing identity recognition and management according to the face feature set; tracking the target to obtain a target person; and recognizing the behavior or emotion of the target person. According to the embodiment of the invention, a complete enforceable scheme from basic target detection, identity recognition, motion trail generation to high-order behavior / emotion analysis and recognition is provided, corresponding optimization is carried out for family nursing tasks / scenes, and the overall enforceability of the scheme and the stability of an algorithm system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent analysis, and more particularly to a method for target perception and behavior analysis. Background Technology

[0002] With the accelerating aging of the population, according to data from the National Bureau of Statistics, the proportion of people aged 60 and above in my country exceeded 20% in 2024. The number of empty-nest elderly households is increasing year by year, and the demand for family care is becoming increasingly urgent. At the same time, the demand for home safety monitoring of children and monitoring of pet activities in dual-income families is also growing. The traditional model of "human care + simple monitoring equipment" is no longer able to meet the diversified and intelligent care needs. Human care has the problem of limited energy and inability to be on duty 24 hours a day. Ordinary monitoring equipment can only realize video recording and playback, lacking active identification and early warning capabilities, and is difficult to deal with emergencies such as falls. From a technological development perspective, while artificial intelligence and computer vision technologies have been widely applied in fields such as public safety and autonomous driving, the unique characteristics of the home environment place higher demands on the technology: the home environment is small and lighting conditions are complex (such as dim lighting at night and backlighting), and existing general-purpose target detection algorithms for network camera terminals suffer from insufficient real-time performance and low recognition accuracy in occluded scenes; in terms of multi-camera collaboration, traditional solutions struggle to maintain cross-device identity consistency, easily leading to confusion in recognizing "multiple identities for the same person"; in the behavior and emotion analysis stage, general-purpose AI models lack sufficient accuracy in recognizing specific behaviors in the home environment (such as elderly people falling or exercising), failing to meet personalized care needs. Against this backdrop, there is an urgent need for an intelligent agent algorithm solution specifically designed for home care scenarios. By integrating technologies such as target detection, identity recognition, trajectory tracking, and behavior analysis, this solution can address core pain points in home scenarios such as multi-task collaboration, adaptability to complex environments, and long-term stable tracking, providing family members with comprehensive and proactive care services. Summary of the Invention

[0003] This invention provides a target perception and behavior analysis method, device, computer equipment, and storage medium to address the urgent need for an intelligent agent algorithm solution specifically designed for home care scenarios. By integrating technologies such as target detection, identity recognition, trajectory tracking, and behavior analysis, it solves core pain points in home scenarios such as multi-task collaboration, adaptability to complex environments, and long-term stable tracking, providing family members with comprehensive and proactive care services.

[0004] A method for goal perception and behavior analysis includes: Detect human bodies and faces in the image, and extract the feature sets of the human bodies and faces; Based on the set of facial features, identity recognition and management are performed; Track the target and obtain information about the target person; Identify the behavior or emotions of the target person.

[0005] Furthermore, the set of human body and facial features includes at least one of the following: human body detection results, facial detection results, ReID features, or facial features.

[0006] Furthermore, the detection of human bodies and faces in the image, and the extraction of the human body and face feature sets, includes: Using human and face detection models, the positions of human bodies and faces in the image are detected in real time, and then human and face features are extracted based on their positions.

[0007] Furthermore, the step of performing identity recognition and management based on the facial feature set includes: Active identity registration and automatic identity addition are performed based on the facial feature set for identity recognition.

[0008] Furthermore, the identity management based on the facial feature set includes a triple management strategy of feature quantity limitation, automatic merging, and user intervention.

[0009] Furthermore, the process of tracking targets to obtain target individuals includes: tracking multiple targets and generating trajectories; and managing the trajectories to obtain the target individuals.

[0010] Furthermore, identifying the target person's behavior or emotions includes using a multimodal large model to analyze and identify the target person's behavior and emotions through single or multiple frames of images.

[0011] A device for target perception and behavior analysis, comprising: The module comprises a detection and extraction module, an identity recognition and management module, a target determination module, and a behavior / emotion recognition module, among which: The detection and extraction module is used to detect human bodies and faces in the image and extract the feature sets of the human bodies and faces; The identity recognition and management module is used to perform identity recognition and management based on the set of facial features; The target determination module is used to track the target and identify the target person; The behavior / emotion recognition module is used to identify the behavior or emotion of the target person.

[0012] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the aforementioned target perception and behavior analysis method.

[0013] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned target perception and behavior analysis method.

[0014] The aforementioned target perception and behavior analysis methods, devices, computer equipment, and storage media provide a complete and feasible solution from basic target detection and identity recognition to motion trajectory generation and advanced behavior / emotion analysis and recognition. All of these solutions have been optimized for home care tasks / scenarios, improving the overall feasibility of the solution and the stability of the algorithm system. Attached Figure Description

[0015] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a schematic diagram of an application environment for a target perception and behavior analysis method according to an embodiment of the present invention; Figure 2 This is a flowchart of a target perception and behavior analysis method according to an embodiment of the present invention; Figure 3 This is a schematic diagram of a target perception and behavior analysis device according to an embodiment of the present invention; Figure 4 This is a schematic diagram of a computer device according to an embodiment of the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] The target perception and behavior analysis method provided in this embodiment of the invention can be applied to, for example... Figure 1The application environment is shown. Specifically, this target perception and behavior analysis method is applied in a target perception and behavior analysis system, which includes, for example, […]. Figure 1 The client and server shown communicate via a network to address the urgent need for an intelligent agent algorithm solution specifically designed for home care scenarios. This solution integrates technologies such as target detection, identity recognition, trajectory tracking, and behavior analysis to solve core pain points in home settings, including multi-task collaboration, adaptability to complex environments, and long-term stable tracking. It aims to provide comprehensive and proactive care services to family members. The client, also known as the user terminal, is the program that provides local services to the client, corresponding to the server. The client can be installed on, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers.

[0019] In one embodiment, such as Figure 2 As shown, a method for target perception and behavior analysis is provided, which is then applied to... Figure 1 Taking the server in the example, the following steps are included: S1. Detect human bodies and faces in the image and extract the human body and face feature sets. The human body and face feature sets include at least one of the following: human body detection results, face detection results, ReID features, or face features. Specifically, a human body and face detection model can be used to detect the positions of human bodies and faces in the image in real time, and then extract human body and face features respectively based on the positions of the human body and face. For example, the human body and face detection models used are both based on lightweight YOLOv11n-Pose fine-tuning. The human body feature extraction model has approximately 600,000 parameters and is also a lightweight pedestrian re-identification network model. The human body detection model is a keypoint detection model, which includes a total of 17 human body keypoints: nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle, and right ankle. By filtering the detection results based on information such as human body bounding boxes, key point locations, and confidence scores, only targets with minimal occlusion are selected. The upper body (excluding arms) of the target is then cropped using key point locations for ReID feature extraction. This effectively reduces feature contamination introduced by occlusion, arm movements, and complex backgrounds.

[0020] The rules for screening human test results are as follows: a. The border confidence score is greater than the threshold th1.

[0021] b. The border width and height are greater than min_size1, where min_size1 is determined based on the image size.

[0022] c. The bounding box confidence is greater than the threshold reid_th, where reid_th ≥ th1.

[0023] d. When the confidence score of a keypoint is less than the threshold kpt_th1, the human body part is considered occluded. Unoccluded keypoints include at least: one head keypoint, one shoulder keypoint, and one hip keypoint.

[0024] If the objective satisfies both conditions a and b, retain it; if the objective satisfies condition c but not condition d, reduce its confidence level. Among the retained targets, those with a final confidence level higher than the preset threshold reid_th are used to extract their ReID features.

[0025] The face detection model is a keypoint-based model, containing a total of five facial keypoints: left eye, right eye, nose, left corner of mouth, and right corner of mouth. This scheme filters the detection results based on information such as the face bounding box, keypoint locations, confidence level, and the Intersection over Union (IOU) between the face and body bounding boxes. Only faces with an IOU greater than a preset threshold and keypoint locations meeting preset conditions are selected for feature extraction. This filters out most blurry and profile targets, improving the accuracy of subsequent face recognition.

[0026] The rules for filtering face detection results are as follows: a. The border confidence score is greater than the threshold th2.

[0027] b. The border width and height are greater than min_size2, where min_size2 is determined by the image size.

[0028] c. The confidence level of any key point is greater than the threshold kpt_th2.

[0029] Faces that do not simultaneously meet all conditions are discarded. Then, the IoU between the body bounding box and the face bounding box is calculated. When the IoU value is less than a preset threshold `iou_th`, the face is considered unrelated to the body. Face bounding boxes are associated with body bounding boxes in descending order of IoU value; faces not associated are discarded. Finally, an affine transformation matrix between the face's keypoints and standard face keypoints is used to perform an affine transformation on the face screenshot to obtain a corrected face. Its features are then extracted for face (identity) recognition.

[0030] The output results include at least one of the following: human body detection results, face detection results, ReID features, or face features. For example, the face feature set (new_feats) contains n vectors of length 512, each representing a feature of n faces.

[0031] S2. Based on the facial feature set, perform identity recognition and management. Specifically, this includes proactive identity registration, automatic identity addition, and identity recognition based on the facial feature set. Specifically, proactive identity registration includes registering identities based on images containing faces, generating FaceIDs for known identities. Automatic identity addition includes generating FaceIDs for unknown identities for faces automatically detected in step S1 during the process. Identity recognition includes using a voting algorithm to match new_feats to existing FaceIDs in the face database. When there are unmatched features in the human body and facial feature set (new_feats), a new FaceID is automatically generated for it; finally, the FaceID for each face is output.

[0032] The identity management system employs a triple management strategy: "feature quantity limit (MAX_FEATURE_NUM) + automatic merging (similarity > face_th1) + user intervention." This strategy avoids redundancy in the identity database and supports data recovery after device restarts, ensuring long-term identity stability. Details are as follows: a. Each time new_feats is passed in, it is merged into the feature set of the corresponding identity. When there are unmatched features in new_feats, a new FaceID is automatically generated for it and stored in the face database.

[0033] b. The face database retains a maximum of MAX_FEATURE_NUM individual face features for each identity. When the number of face features for a certain identity is greater than MAX_FEATURE_NUM (e.g., MAX_FEATURE_NUM + 1), the face feature with the highest internal cosine similarity is removed.

[0034] c. The algorithm periodically calculates and compares the average cosine similarity of faces among all identities, automatically merging identities with a similarity higher than the threshold face_th1 into one identity, and then informing the tracking module of the merged identity. Users can also initiate the merging of certain identities.

[0035] d. Store identity information on disk, and delete identity information that is no longer maintained from disk so that the algorithm can be restored when the device is restarted.

[0036] The "voting system" mentioned in identity recognition is designed to improve recognition accuracy. Compared to the traditional "single feature comparison" scheme, its anti-interference ability is significantly enhanced, specifically: 1. Calculate the cosine similarity matrix S-matrix between new_feats and all face features history_feats for all identities in the face database. Mark each position in the matrix with a value greater than the threshold face_th2 as 1 vote.

[0037] 2. Count the number of votes each identity gives to new_feats Ultimately, each identity gives a vote score to new_feats. Where m is the number of identities in the face database, and n is the number of features in new_feats. This represents the total number of votes given by the i-th identity in the face database to the j-th feature in new_feats. The number of facial features for the i-th identity. If the value is less than the threshold score_th, the j-th feature in new_feats is considered not to match the i-th identity. Set to 0.

[0038] 3. Select the elements in P that are greater than 0, and match them with the identities of new_feats in descending order of their values.

[0039] S3. Track the target to obtain the target person. Specifically, this includes: tracking multiple targets and generating trajectories; managing the trajectories to obtain the target person. For example, the trajectory generation includes: using an improved BotSORT algorithm to generate the motion trajectories of multiple targets, which facilitates subsequent behavior analysis.

[0040] The trajectory management includes: a. Store the trajectories of known identities on disk, and delete the trajectories that are no longer maintained from disk, so that the algorithm can be restored when the device is restarted later; b. Tracks with high similarity in human appearance characteristics are considered to belong to the same person and their tracks are merged; the similarity of human characteristics in tracks under maintenance is checked regularly. c. Regularly clean up unknown identity traces that have not been updated for a long time.

[0041] The target tracking algorithm includes the following improvements based on BotSORT: Improvement 1: Three-stage matching Phase 1: Matching trajectories for identified targets. The identified targets are directly matched one-to-one with their FaceID and historical trajectory. Since the FaceID of the person has already been obtained through face detection and recognition in the M1 module, matching using FaceID greatly improves the matching accuracy in complex scenes. The remaining unmatched historical trajectories proceed to the second stage.

[0042] Phase Two: Matching Trajectories for High-Confidence Targets This scheme designates targets with a confidence level higher than reid_th as high-confidence targets, and the remaining targets as low-confidence targets.

[0043] High confidence levels include ReID feature information; therefore, when calculating the cost matrix, the IoU distance matrix of the human bounding box and the cosine distance matrix of the ReID features are fused. Let... Let be the cosine distance between the ReID features of the i-th trajectory and the ReID features of the j-th detection. Let be the IoU distance between the predicted bounding box of the i-th trajectory and the bounding box of the j-th detected human body. These are elements of cost matrix C. Then, the Jonker-Volgenant algorithm is used to solve the matching problem between the detected target and the remaining historical trajectories. The remaining unmatched historical trajectories proceed to the third stage.

[0044] Phase 3: Matching trajectories for targets with low confidence levels. Low-confidence targets do not contain ReID feature information, so the IoU distance matrix is ​​directly used as the cost matrix. Then, the Jonker-Volgenant algorithm is used to solve the matching problem between low-confidence targets and remaining historical trajectories.

[0045] Improvement point 2: Identity / Trajectory Merging During the long-term operation of the algorithm, many trajectories are generated as people appear and disappear, including trajectories of different people and trajectories of the same person. However, not every trajectory can identify a face during its generation. In order to analyze the behavior of people over a relatively long period of time, we need to assign an identity to the trajectories of unknown identities as much as possible. This solution uses an identity / trajectory merging method to achieve this requirement.

[0046] 1. When two targets appear simultaneously, their corresponding trajectories should be two different identities. Therefore, during the tracking process, it is necessary to record which trajectories do not belong to the same identity and cannot be merged.

[0047] 2. A certain number of ReID features also need to be cached during the tracking process, with the latest 50 stored by default.

[0048] 3. Calculate the average cosine similarity of the ReID features between each pair of maintained trajectories. If there are trajectories with a similarity higher than the threshold merge_th and not within the range that cannot be merged, then merge the two trajectories.

[0049] S4: Identify the target person's behavior or emotions. Specifically, this includes using a multimodal large model to analyze and identify the target person's behavior and emotions through single or multiple frames of images, including: eating, falling down, exercising, entering or leaving a room, being happy, being calm, etc.

[0050] For example, the following steps may be included: 1. Based on the movement trajectory of the target person, extract a video image sequence Seq0 within a certain time range, and sample and extract a fixed number of image frames Seq1. Mark the identity and location information of the target person in the extracted image sequence Seq1 to obtain Seq2.

[0051] 2. Input the Seq2 algorithm along with the designed prompts into the fine-tuned multimodal large model to obtain the output.

[0052] For different caregivers, the prompts can be flexibly adjusted to guide the model to output different behavioral recognition results. For example, it can focus on whether the elderly have fallen or whether children have come into contact with dangerous items.

[0053] This invention provides a complete and feasible solution for home care scenarios, encompassing basic object detection, identity recognition, motion trajectory generation, and advanced behavior / emotion analysis and recognition. Each solution is optimized for home care tasks / scenarios, improving overall feasibility and algorithm stability. The invention employs a lightweight network model with multi-threaded processing to optimize NPU utilization and enhance overall inference speed. Its modular design facilitates flexible deployment. The "layered perception + multimodal fusion" architecture decouples the functions of each module while also enabling reasonable data fusion between them. It allows for flexible selection based on family needs (e.g., enabling S1 and S3 for basic monitoring and adding S4 for emotion analysis). Key point filtering and feature extraction optimization improve ReID feature stability. Combined with identity merging (automatic cosine similarity merging) and trajectory management (timely cleaning of invalid trajectories and persistent disk storage), it achieves long-term stable tracking, solving the "identity confusion" problem and preventing identity loss due to camera switching or temporary obstruction, thus reducing the error rate. Based on the unique identity identifier FaceID, it associates trajectories belonging to the same identity across channels, maintaining consistency in the identity of individuals across channels and making the trajectory of the monitored individual more complete. Employing a large multimodal model to identify human behavior / emotions, supported by its massive pre-training data and large parameter set, we only need to fine-tune with a small amount of data to achieve results exceeding those of traditional network models, while also facilitating future expansion to recognize different customized behaviors.

[0054] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0055] In one embodiment, a target perception and behavior analysis device is provided, which corresponds one-to-one with the target perception and behavior analysis methods described in the above embodiments. For example... Figure 3 As shown, the target perception and behavior analysis device 100 includes a detection and extraction module 101, an identity recognition and management module 102, a target determination module 103, and a behavior / emotion recognition module 104. Detailed descriptions of each functional module are as follows: The detection and extraction module 101 is used to detect human bodies and faces in the image and extract the human body and face feature sets; The identity recognition and management module 102 is used to perform identity recognition and management based on the set of facial features; The target determination module 103 is used to track the target and obtain the target person; The behavior / emotion recognition module 104 is used to identify the behavior or emotion of the target person.

[0056] Specific limitations regarding the target perception and behavior analysis device can be found in the limitations of the target perception and behavior analysis method described above, and will not be repeated here. Each module in the aforementioned target perception and behavior analysis device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0057] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a target perception and behavior analysis method.

[0058] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the target perception and behavior analysis method described in the above embodiments, for example... Figure 2S1-S4, as shown, will not be described again here to avoid repetition. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in this embodiment of the target perception and behavior analysis device, for example... Figure 3 The functions of the detection and extraction module 101 shown are not described again here to avoid repetition.

[0059] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When executed by a processor, the computer program implements the target perception and behavior analysis method described in the above embodiments, for example... Figure 2 S1-S4, as shown, will not be repeated here to avoid repetition. Alternatively, when the computer program is executed by the processor, it implements the functions of each module / unit in this embodiment of the target perception and behavior analysis device, for example... Figure 3 The functions of the detection and extraction module 101 shown are not described again here to avoid repetition. The computer-readable storage medium can be non-volatile or volatile.

[0060] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0061] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0062] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for target perception and behavior analysis, characterized in that, include: Detect human bodies and faces in the image, and extract the feature sets of the human bodies and faces; Based on the set of facial features, identity recognition and management are performed; Track the target to obtain information about the target person; Identify the behavior or emotions of the target person.

2. The method for target perception and behavior analysis as described in claim 1, characterized in that, The set of human body and facial features includes at least one of the following: human body detection results, facial detection results, ReID features, or facial features.

3. The method for target perception and behavior analysis as described in claim 2, characterized in that, The human body and face in the detected image, and the extraction of the human body and face feature set include: Using human and face detection models, the positions of human bodies and faces in the image are detected in real time, and then human and face features are extracted based on their positions.

4. The method for target perception and behavior analysis as described in claim 1, characterized in that, The process of identity recognition and management based on the facial feature set includes: Active identity registration and automatic identity addition are performed based on the facial feature set for identity recognition.

5. The method for target perception and behavior analysis as described in claim 1, characterized in that, The identity management based on the facial feature set includes a triple management strategy: feature quantity limitation, automatic merging, and user intervention.

6. The method for target perception and behavior analysis as described in claim 1, characterized in that, The process of tracking targets to obtain target individuals includes: tracking multiple targets and generating trajectories; and managing the trajectories to obtain the target individuals.

7. The method for target perception and behavior analysis as described in claim 1, characterized in that, The identification of the target person's behavior or emotions includes using a multimodal large model to analyze and identify the target person's behavior and emotions through single or multiple frames of images.

8. A device for target perception and behavior analysis, characterized in that, include: The module comprises a detection and extraction module, an identity recognition and management module, a target determination module, and a behavior / emotion recognition module, among which: The detection and extraction module is used to detect human bodies and faces in the image and extract the feature sets of the human bodies and faces; The identity recognition and management module is used to perform identity recognition and management based on the set of facial features; The target determination module is used to track the target and identify the target person; The behavior / emotion recognition module is used to identify the behavior or emotion of the target person.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the target perception and behavior analysis method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the target perception and behavior analysis method as described in any one of claims 1 to 7.