Community personnel multi-modal recognition system and method based on image perception large model
By constructing a role priority matrix and dynamically optimizing it based on a large image perception model for community personnel multimodal recognition, the problem of low recognition efficiency in existing technologies is solved, and efficient and accurate recognition is achieved in complex scenarios.
Patent Information
- Application Number
- CN202511133798.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-08-14
AI Technical Summary
Existing community personnel identification systems are inefficient in complex or flexible scenarios and cannot automatically adjust identification priorities according to specific scenarios and identification accuracy, resulting in identification deviations and response delays.
A multimodal identification method for community members based on a large image perception model is adopted. By parsing the community event labels of image frames, a role priority matrix is constructed, the bounding boxes of target individuals are extracted and instances are segmented, behavioral fragments are extracted based on optical flow features and dynamic trajectories, role assignment and face recognition sequence are determined, and the role priority matrix is dynamically optimized through a reinforcement learning strategy network.
It improves the recognition efficiency and accuracy in complex behavioral interaction scenarios, reduces the recognition failure rate and response latency, and adapts to the complex behavioral interaction relationships in high-density community scenarios.
Smart Images

Figure CN120726674B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image recognition, more particularly, to a community personnel multi-modal recognition system and method based on an image perception large model. BACKGROUND
[0002] In the community image recognition scene of multi-target dense co-occurrence, the system needs to complete the identity recognition task of multiple individuals within a limited time, and the recognition efficiency is highly dependent on the setting of the recognition order, and the priority judgment of the main target becomes a key link that affects the accuracy of the recognition result and the overall response speed. Most of the existing systems use fixed role rules or visual features such as image saliency, target motion amplitude, image center offset to sort individuals. This kind of processing method can still play a role in scenes where the personnel distribution is relatively simple or the behavior structure is clear, but in actual scenes such as kindergarten pick-up, elderly care, and visitor gathering, where the behavior density is high and the semantic role relationship is complex, recognition deviation frequently occurs. The system often takes individuals with fixed identity priority as the main recognition object, and cannot automatically adjust the recognition priority according to the specific scene and recognition accuracy, resulting in low recognition efficiency, permission release errors or response delays in some complex or flexible scenes. SUMMARY
[0003] In order to overcome the above-mentioned defects of the prior art, the embodiments of the present application provide a community personnel multi-modal recognition method based on an image perception large model to solve the problems raised in the background art.
[0004] To achieve the above object, the present application provides the following technical scheme:
[0005] The community personnel multi-modal recognition method based on an image perception large model comprises the following steps:
[0006] S1: parsing the community event label of the image frame from the community monitoring image, and presetting the role priority matrix driven by the event;
[0007] S2: extracting all detectable target individuals in the continuous frames of the monitoring image, performing boundary box positioning and instance segmentation on the target individuals, and constructing a multi-target image entity set;
[0008] S3: inputting the target individual structure in the target image entity set into a time sequence action recognition network, extracting the behavior segment of the target individual based on the optical flow feature and the dynamic trajectory, and performing action category labeling;
[0009] S4: establishing action template labels of different roles, and according to the semantic matching of the target individual action label and the action template label, assigning roles to the target individual and calibrating the face recognition order;
[0010] S5: performing face recognition scheduling in the face recognition period based on the face recognition sequence calibration result, and calculating a face recognition success rate;
[0011] S6: constructing a reinforcement learning strategy network with the role priority matrix as a state space, and updating the weight values of the role priority matrix.
[0012] In a preferred embodiment, in S1, the community event label of the image frame is parsed from the community monitoring image, and the preset event-driven role priority matrix specifically includes:
[0013] The timestamp and image shooting location of the image frame are extracted from the community monitoring image as event binding conditions;
[0014] The event type label is preset according to the regional nature of the community, the community event scheduling log is retrieved based on the event binding conditions, and the event type label to which the current image belongs is identified;
[0015] According to the event type label, the identification target role list of the event and the corresponding identification priority are set, and the role priority weight value set under the context of the current image frame is generated;
[0016] Based on the association structure of the role list and the role priority weight value set, the role priority matrix corresponding to different event labels is constructed.
[0017] In a preferred embodiment, in S2, all detectable target individuals in the continuous frames of the monitoring image are extracted, the boundary box positioning and instance segmentation are performed on the target individuals, and the multi-target image entity set is constructed specifically includes:
[0018] The pixel matrix size alignment is performed on the community monitoring image frame sequence, and all image frame data is converted into a standardized single color channel format;
[0019] The standardized image frame sequence is input into the deep target recognition network model, and the boundary box of all target individuals is output. The mask generation is performed on the boundary box area, and the pixel mask graph matched with the target individual contour is extracted;
[0020] The corresponding target individual boundary box area, mask graph data and time index in the standardized image frame are integrated into the target individual structure;
[0021] All target individual structures are organized in chronological order to form a multi-target image entity set across frames.
[0022] In a preferred embodiment, in S3, the target individual structure in the target image entity set is input into the time sequence action recognition network, the behavior segment of the target individual is extracted based on the optical flow feature and dynamic trajectory, and the action category label is specifically includes:
[0023] Interframe cropping is performed on the bounding box regions of each target individual in the multi-target image entity set to construct an image frame sequence of the corresponding target individual;
[0024] Interframe optical flow calculation is performed on the target individual image frame sequence, and an optical flow tensor feature composed of pixel-level motion vectors is generated;
[0025] The target individual image frame sequence is spliced into a bounding box trajectory, and a short-time behavior feature of the target individual is constructed based on the joint of the optical flow tensor feature and the bounding box trajectory;
[0026] The short-time behavior feature is input into a time sequence action recognition network, and an action class probability distribution is output. The action label with the maximum probability is written back to the target individual structure corresponding to the multi-target image entity set.
[0027] In a preferred embodiment, in S4, the action template label of different roles is established, and the role assignment and face recognition order calibration of the target individual are performed according to the semantic matching of the target individual action label and the action template label, which specifically includes:
[0028] Obtain the role list corresponding to the current event label, and construct a role action template;
[0029] Perform semantic embedding processing on the action label of the target individual and the role action template to generate action vectors and role action template vectors that match in the same semantic space;
[0030] Vector similarity calculation is performed on the action vectors and role action template vectors in the semantic space, and a role matching degree matrix is output according to the calculation result;
[0031] The role label corresponding to the maximum matching degree is taken as the role assignment of the target individual in the current event context, and is written into the corresponding target individual structure;
[0032] Based on the matrix element product of the role matching degree matrix and the role priority matrix, the target individual is marked in face recognition priority order according to the descending order of the result.
[0033] In a preferred embodiment, in S5, face recognition scheduling is performed in the face recognition period based on the face recognition order calibration result, and the face recognition success rate is calculated, which specifically includes:
[0034] Based on the real-time obtained community monitoring image, the face recognition period with variable length is divided according to the number of detectable target individuals in the image, and the target individual in the monitoring image is scheduled for face recognition according to the face recognition order calibration;
[0035] Wherein, each face recognition period takes the first successful recognition as the end marker, and the face recognition success rate is calculated based on the total number of face recognition attempts and the number of successful attempts in each face recognition period.
[0036] In a preferred embodiment, in S6, the reinforcement learning strategy network is constructed with the role priority matrix and the role matching degree matrix as the state space, and updating the weight values of the role priority matrix specifically includes:
[0037] The role matching degree matrix and the role priority matrix generated in the current face recognition period are collected, spliced and constructed into a state representation tensor in the form of column vectors, and used as the state input of the strategy network;
[0038] The face recognition success rate index in the current face recognition period is obtained, the role label that successfully recognizes is obtained based on the proportion of the role label, and the recognition success rate index is used as the reward value of the corresponding role label;
[0039] A deep reinforcement learning model based on the change of the strategy gradient is established, the state tensor is used as the strategy input, and the role label reward value is used as the reward, and an update strategy function of the role priority weight is constructed;
[0040] A network parameter optimization process based on gradient descent is performed, and the weight matrix and the bias vector of the multi-layer perception structure of the strategy function are trained and updated by back propagation round by round;
[0041] After the strategy network converges, the role priority weight update result under the current state is output, and the original role priority matrix is replaced.
[0042] On the other hand, the present application provides a community personnel multi-modal recognition system based on an image perception large model, which comprises a role priority construction module, a target extraction module, an action labeling module, a role assignment module and a role priority adjustment module, wherein:
[0043] The role priority construction module: according to the timestamp and the shooting position of the image frame, the community event scheduling log is searched to determine the event label, and the recognition target role list and the priority weight are set according to the event label, and the role priority matrix corresponding to the current image frame is constructed;
[0044] The target extraction module: according to the community monitoring image frame sequence, the size standardization processing is performed, the boundary box and the mask information of all detectable target individuals are extracted by inputting the deep target recognition network model, and a multi-target image entity set continuous across frames is constructed;
[0045] The action labeling module: the image sequence of the individuals in the target image entity set is cut and the optical flow tensor is calculated between frames, the short-time behavior features are constructed combined with the boundary box track, and the action class label is generated by inputting into the time sequence action recognition network;
[0046] The role assignment module: based on the event role list, an action template is constructed, semantic embedding and template matching are performed on the target individual action label, a role matching degree matrix is generated, and the target is sequentially sorted by face recognition according to the product of the matching degree and the role priority matrix;
[0047] The role priority adjustment module: the recognition success rate of the identification cycle is counted, the state input tensor is constructed based on the role matching degree matrix and the role priority matrix, the reward signal is generated combined with the recognition success rate, the strategy gradient type reinforcement learning network is trained, and the role priority weight adjustment result is output to update the role priority matrix.
[0048] The technical effects and advantages of the community personnel multi-modal recognition system and method based on the image perception large model of the present application are as follows:
[0049] By introducing the role list and the event driven matrix construction mechanism, the system recognition strategy can respond to the semantic difference of the scene, and the spatial layout and behavior semantics are considered. The behavior information of the target individual is extracted by using optical flow and dynamic trajectory, and semantic level role reasoning is realized combined with the action template, which effectively enhances the recognition ability of the dominant role in the complex behavior interaction scene. In the identification scheduling process, the face recognition success rate is collected in real time, the role priority matrix is dynamically optimized through the reinforcement learning strategy network, so that the identification sorting mechanism has self-adaptive adjustment ability, and the overall recognition efficiency and accuracy are effectively improved. It can adapt to the complex behavior interaction relationship in the high-density community scene, significantly improve the focusing performance of the system on the key identification object, and reduce the recognition failure rate and response time delay. BRIEF DESCRIPTION OF DRAWINGS
[0050] Figure 1 The present application is a community personnel multi-modal recognition method based on an image perception large model.
[0051] Figure 2 The present application is a community personnel multi-modal recognition system based on an image perception large model. DETAILED DESCRIPTION
[0052] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0053] Embodiment 1, Figure 1 The present application is a community personnel multi-modal recognition method based on an image perception large model, which comprises the following steps:
[0054] S1: Analyze the community event label of the image frame from the community monitoring image, and preset the event-driven role priority matrix;
[0055] S2: Extract all detectable target individuals in the continuous frames of the monitoring image, perform boundary box positioning and instance segmentation on the target individuals, and construct a multi-target image entity set;
[0056] S3: Input the target individual structure in the target image entity set into a time sequence action recognition network, extract the behavior segment of the target individual based on the optical flow feature and dynamic trajectory, and perform action category labeling;
[0057] S4: Establish action template labels for different roles, assign roles to target individuals and calibrate face recognition order according to the semantic matching of target individual action labels and action template labels;
[0058] S5: Perform face recognition scheduling in the face recognition cycle based on the face recognition order calibration result, and calculate the face recognition success rate;
[0059] S6: Construct a reinforcement learning strategy network with the role priority matrix as the state space, and update the weight value of the role priority matrix.
[0060] In S1, the community event label of the image frame is analyzed from the community monitoring image, and the event-driven role priority matrix is preset.
[0061] The video stream from the community monitoring system is decoded frame by frame, and the basic attribute information corresponding to each frame image is extracted from it. The basic attribute information includes but is not limited to the timestamp information of the image frame (i.e. the specific date and time of the image frame acquisition) and the image shooting position identifier (i.e. the device number and installation area number of the corresponding camera). Among them, the installation area of the camera needs to be location coded in the community monitoring system in advance, for example, the kindergarten entrance area is marked as "E-G01", the old people's apartment passage is marked as "E-A02", and the entrance registration front desk area is marked as "E-R01", etc., to ensure that the event semantic context can be quickly established according to the area number in the subsequent processing steps. According to the functional area attribute corresponding to the camera installation position, the system reads the area preset event type label from the configuration database. For example, if the camera is installed in the "E-G01" area, this area is marked as "child pick-up area", and its corresponding preset event type label is "parent pick-up"; if the camera is located in the "E-A02" area, the area event type label is "old people passing" or "accompanying passing"; if it is located in the "E-R01" area, it may be "visitor registration" and the like. Based on the timestamp and shooting position of the image frame, the system takes the above two information as the event binding condition, and searches in the community dispatch management platform whether there is a dispatch record at this time point and in this area, such as whether there is a collective school dismissal, nurse shift change, visitor registration arrangement, etc. If the log record exists, the event description in the record is taken as the event type label of the current image frame; if there is no clear record, the default area attribute label is used as the event type of the image.
[0062] After identifying the event type label, the system will enter the role priority setting process. Each type of event type label corresponds to a set of pre-defined identification target role list and initial identification priority. For example, in the "parent pick-up" event, the role list may include "children", "parents", and "teaching staff", and the initial priority is set as "teaching staff" 1.0, "parents" 0.8, and "children" 0.4; in the "old people passing" event, the role list is "old people", "accompanying personnel", and "security personnel", and the priority is "old people" 1.0, "accompanying personnel" 0.6, and "security personnel" 0.3; the initial value of the priority is set by the system administrator or the community operator according to business needs or historical identification data statistics. After the role priority weight value set is generated, the system constructs the role priority matrix of the current image frame according to the role order of the set. The matrix takes the role list as the column vector, and the corresponding role priority weight value forms a one-dimensional structure, which is used as the prior strategy input of the face recognition task in the event scene of the image frame, for identification sorting or weight regulation in the subsequent identification task. If there are multiple image frames belonging to the same event type in actual deployment, the system will construct a set of fixed format role priority matrices for each type of event and use them uniformly.
[0063] In S2, all detectable target individuals in the monitoring image continuous frames are extracted, the boundary box positioning and instance segmentation are performed on the target individuals, and the multi-target image entity set is constructed.
[0064] The unified image preprocessing operation is performed on the original image frame sequence continuously collected by the community monitoring system to ensure the consistency of the input data of the subsequent multi-target recognition model. The original image frames may come from different brands or models of camera devices, and there are inconsistencies in resolution, channel number, image format, etc. Therefore, size alignment and color channel uniform processing need to be completed in the data preprocessing stage. The size alignment operation resamples all image frames to a fixed resolution, which is set to 1280x720 pixels in this embodiment. This resolution is selected mainly considering the premise of ensuring the distinguishability of target details and taking into account the model processing efficiency. If the size of the original image frame is greater than this value, scaling is performed and then the image is centered and cropped. If it is smaller than this value, edge padding is performed and then scaling is performed. This process uses linear interpolation to avoid the occurrence of jagged or image blur phenomena. Simplify the model input and highlight the outline structure information, and convert the color image to a single color channel format. In this embodiment, the gray channel processing method is used, and the RGB three channels are fused into a single channel image through the weighted average formula, and the specific weight parameters are set as red channel 0.2989, green channel 0.5870, and blue channel 0.1140.
[0065] After the image frame is standardized, it is input into the YOLOv7 deep target recognition network model for target detection. The YOLOv7 model has high precision and high real-time performance, and is suitable for detecting dense targets in community monitoring images. The model has been pre-trained based on the COCO dataset and fine-tuned and optimized with the community dataset to ensure high recognition rate for common roles such as the elderly, children, visitors, property personnel, etc. In the model detection stage, YOLOv7 outputs the bounding box coordinates and confidence scores of all detected targets in each frame of image. To improve the accuracy of data processing, only the bounding box results with a confidence score greater than 0.5 are retained. This threshold is determined by experimental verification to effectively filter out false positives while ensuring recognition rate. For the above output bounding box region, a pixel-level mask generation operation is further performed. This embodiment uses an instance segmentation-based boundary mask method, which inputs the image region within the bounding box into a feature extractor and calls a clustering segmentation algorithm to generate a mask map matching the target individual contour. The mask map is represented in binary form: a pixel point with a mask value of 1 represents the valid region belonging to the detected target, and a pixel point with a mask value of 0 is the background region. This processing step is particularly crucial for separating targets that overlap or have complex contours in the image, and can avoid boundary confusion in subsequent behavior feature analysis. After completing the extraction of bounding boxes and masks, the target individual structures extracted from each frame of image are organized and encapsulated. The target individual structure should include the following fields: bounding box coordinate data (for image cropping), mask map data (for shape analysis), image frame time index (for time sequence splicing), and camera number and location identifier of the target image frame (for spatial positioning).
[0066] The above structure is arranged in the order of image frame time and aggregated to form a multi-target image entity set that is continuous across frames. This set will serve as the basic data structure for subsequent action recognition and role assignment, supporting the transition from single-frame recognition to multi-frame behavior understanding. In practical applications, the number of image frames per second can be set to 15 frames. The system processes 1 second of image every second, generating an image entity set containing 15 frames of image and an average of 20 target individual structures, ensuring that the recognition processing needs of dense scenes are met.
[0067] In S3, the target individual structures in the target image entity set are input into the time sequence action recognition network to extract the behavior segments of the target individuals based on optical flow features and dynamic trajectories and label the action categories.
[0068] Based on the multi-target image entity set constructed in the previous step, the frame data cutting process is performed on each target individual structure in the set. This operation takes the bounding box coordinates of the target individual in the image frame sequence as the cutting reference, and extracts the image content of the bounding box region frame by frame according to its time index, and generates the image frame sequence corresponding to the target individual. In the image cutting process, to ensure spatial information consistency, the size of the bounding box extracted in all frames of each target individual is uniformly set to the maximum bounding box size as the standard, and the smaller frame bounding box is filled by edge extension to avoid introducing discrimination interference due to target scale changes. In terms of time length, to construct a short action sequence, the system sets the frame sequence length to 10 frames, the step size to 1 frame, and the time interval to about 0.66 seconds (based on 15 frames per second). This parameter value can be adjusted according to the frequency of dynamic behavior in the scene.
[0069] After cutting, the system performs inter-frame optical flow calculation on each target individual image frame sequence, aiming to obtain the pixel-level motion information of the target between consecutive frames. The optical flow calculation uses the classic dense optical flow method, which can output the motion vector of each pixel point between two frames. To enhance robustness in complex occlusion, blur, or low-contrast areas, the image is first subjected to contrast enhancement and noise suppression processing before optical flow calculation. The output result is stored in the form of an optical flow tensor, with tensor dimensions including image height, width, and two-dimensional motion vector (horizontal and vertical components). Each target individual sequence will generate multiple optical flow tensors, forming the continuous motion features of the individual within the current time window. While generating the optical flow tensor features, the image frame sequence of the target individual is spliced in time order to form a bounding box trajectory. The bounding box trajectory is a time sequence description of the position of the target individual in each frame, which can be regarded as a spatial movement path in time sequence form. This path not only retains the displacement information of the target individual in the image, but also provides key clues for subsequent judgment of behavior rhythm, direction transformation, and speed change. The trajectory data includes properties such as center point coordinates, bounding box size, and frame index, and can be used to construct a smooth trajectory curve through interpolation.
[0070] After completing the construction of optical flow tensors and bounding box trajectories, the system jointly encodes these two types of features into short-term behavior features. The behavior feature construction process includes three parts: first, spatial down-sampling and channel normalization are performed on the optical flow tensor to compress the calculation amount and retain the main motion information; second, polynomial fitting and time normalization are performed on the bounding box trajectory to extract trajectory differential features such as speed and direction change rate; finally, the motion information and path information are integrated into a unified behavior feature vector through feature splicing. This vector sequence serves as the dynamic behavior representation of the target individual within a short time window.
[0071] The behavior feature vector sequence constructed above is input into a pre-trained time sequence action recognition network. The recognition network used can be constructed based on a time sequence convolution network or a bidirectional gated recurrent network structure, and its training data is derived from community behavior collection, covering typical action categories such as hand-holding, stopping, running, bending, raising hands, and pointing, and the network output is a probability distribution vector of the action category. The system selects the action category label with the highest probability from the distribution as the recognition result and writes it into the "action label" field of the corresponding target individual structure in the multi-target image entity set.
[0072] In S4, action template labels of different roles are established, and the target individual is assigned a role and the face recognition order is calibrated according to the semantic matching of the target individual action label and the action template label.
[0073] When the system completes the action category labeling of the target individual and obtains the event label of the current image frame, it first enters the role list extraction and action template construction step. The system presets the main event types in the community monitoring scene and their corresponding role lists, and all event type labels are statically configured in the community platform database. Taking the "parent pick-up" event at the kindergarten gate as an example, the role list corresponding to this event label includes "children", "parents", and "staff". Under each role item, the system pre-configures typical action templates, such as the "children" role containing "running", "looking around", "jumping", etc. action labels, the "parents" role containing "waving hands", "holding hands", "carrying objects", etc. action labels, and the "staff" role containing "pointing", "checking", "taking pictures", etc. action labels. The construction of action templates uses a combination of human business experience and historical monitoring data, and administrators can maintain the action label list in the system background. The system regularly calculates the action recognition frequency in historical monitoring data to dynamically modify and expand the templates. In addition, to ensure the diversity of matching, the template action label allows multiple label configuration, each role can correspond to several typical action templates, and the action label can contain sub-labels of the same action in different scenarios (such as "waving hands-greeting" and "waving hands-signaling"). The execution of this step ensures that the subsequent semantic matching has the support of role behavior diversity, ensuring flexible adaptation in different scenarios.
[0074] For each action label stored in the target individual structure, semantic embedding processing is performed for vectorized matching with the role action template. Specifically, the system preinstalls a natural language processing model to convert the action label into a vectorized semantic representation, which is specially fine-tuned based on a community dataset and adapted to the action label system in the monitoring scenario. The semantic embedding method adopted is Word2Vec or BERT pre-training model, and the action label is uniformly mapped into a semantic vector space through average pooling of word vectors or sentence vector generation mechanism. In the semantic embedding processing, to ensure the stability of label conversion, the system performs pre-cleaning on the input action label, such as uniform word form, symbol removal, and correction of misspelled words. For the role action template label, the system also performs the same semantic embedding process to ensure that the vector representations of the action label and the role action template are in the same semantic space, thereby facilitating subsequent vectorized similarity calculation. Through this step, the system not only completes the vectorized representation of the label, but also lays a data foundation for subsequent role matching matrix construction, ensuring the semantic comparability and computational uniformity among multiple role action labels. The matching degree between the action label and each role action template is calculated based on the semantic vector. Specifically, the system calculates the vector similarity between the action vector of each target individual and the action template vector of each role in the role list one by one. The similarity calculation method uses cosine similarity or the inverse of Euclidean distance. The administrator can set a threshold value in the background, such as a minimum similarity threshold of 0.7. Matching pairs below this threshold are marked as weak matching relationships. Each calculation result forms a role matching score. The system arranges all target individual matching scores for all roles into a matrix form (role matching matrix), with the matrix row representing the target individual, the matrix column representing the role list, and the matrix element being the similarity score of the target individual and the role action template. The maximum value in the role matching matrix and its corresponding role label are extracted, and this label is taken as the final role assignment of the target individual in the current event context.
[0075] The role semantics and scene priority are integrated into one, and the system first performs matrix element multiplication operation on the role matching matrix and the role priority matrix in the current event scene to generate a role priority score matrix. The role priority matrix is derived from the priority weight value set in the event label step. The result of matrix element multiplication represents the action semantic matching degree and scene role priority weight that each target individual has in the current scene. The system then performs descending order sorting on the priority score of each target individual to generate a face recognition order list of the target individual, and writes the order designation into the target individual structure.
[0076] In S5, face recognition scheduling is performed based on the face recognition order designation result within the face recognition cycle, and the face recognition success rate is calculated.
[0077] Continuous image frame data is acquired from a real-time image stream continuously collected by the community monitoring device, and multi-target detection and segmentation processing is performed on each image frame to extract all detectable target individual structures. To adapt to the characteristics of high mobility of personnel in a dynamic environment, the "face recognition period" is specially designed in this step, and a fixed time window is not used to divide the period, but the period length is dynamically divided according to the number of detectable target individuals in the current image. Specifically, the system counts the number of detectable target individuals in the current frame and its adjacent frames in the sequence of continuous frames according to the target detection result. If the number of targets in consecutive frames changes by more than a set threshold (such as ±2), the system marks the frame as the starting point of a new recognition period, ensuring that the number of target individuals and their distribution are relatively stable within each face recognition period. The threshold is set through system testing, which can reflect the significant changes in crowd density and avoid excessively short periods caused by excessive division. The end condition of the period is initially set to automatically switch to a new period after the next target number change exceeds the threshold or the fixed maximum period length (e.g., 5 seconds) is reached. Through this dynamic division based on the number of targets, the stability of dense crowd flow and recognition scheduling is effectively balanced.
[0078] After determining the starting point of the current face recognition period, the target individual face recognition order labeling result output in the previous step is first retrieved. This order labeling is generated by the product of the role matching degree and the role priority weight matrix, representing the face recognition priority of the target individual in the current event scene. The system queues all target individuals in the current period in order of priority based on this order labeling, forming a recognition scheduling queue. The individuals in the queue are scheduled to the face recognition engine one by one for recognition attempts. After each recognition attempt, the system records the attempt result in real time and marks the target individual who has been successfully recognized as "recognition completed". If a target individual is recognized during the recognition process, it will not be re-identified in the same period to avoid wasting resources. If all individuals in the queue are recognized or the period ends, the system archives the queue results and starts a new round of recognition scheduling for the period. During the scheduling process, the system can dynamically adjust the scheduling pace according to the priority of the target individuals in the queue, such as setting a shorter recognition waiting time or increasing the number of attempts for high-priority targets. The specific parameters can be pre-set according to the actual use requirements of the community, such as setting the waiting time for high-priority targets to 200ms and the waiting time for low-priority targets to 500ms.
[0079] In addition, the end mark of each face recognition cycle is also set through the first successful recognition event. Specifically, when the system detects that any target individual in the queue successfully passes the face recognition verification during the face recognition scheduling process, the system immediately records the current successfully recognized target information and timestamp, and takes this event as the end mark of the current cycle. The purpose of this is to maximize the shortening of the cycle length on the basis of ensuring the rapid passage of key targets, thereby improving the overall recognition response speed. For each face recognition cycle, the system will record the total number of face recognition attempts of all target individuals in the cycle, and count the number of successful recognitions. The total number of attempts refers to the number of recognition calls of the system to all target individuals, and the number of successes refers to the number of times the system obtains valid identity information of a target individual in one call. The system calculates the recognition success rate indicator of the current face recognition cycle by taking the ratio of the number of successes to the total number of attempts. For example, if the total number of attempts in a certain cycle is 15 and the number of successes is 3, then the recognition success rate of the cycle is 0.2 (i.e., 20%). The system takes this indicator as a quantitative feedback of the performance of the face recognition in the current cycle, records it in the recognition log database, and writes it into the strategy update module together with the target role matching degree matrix and the role priority matrix corresponding to the current cycle, to provide data support for the dynamic adjustment of the role priority weight in the subsequent cycle.
[0080] In S6, a reinforcement learning strategy network is constructed with the role priority matrix and the role matching degree matrix as the state space, and the weight value of the role priority matrix is updated.
[0081] After the system completes the recognition scheduling and recognition success rate statistics of the current face recognition cycle, the role matching degree matrix and the role priority matrix generated in the current cycle are first collected and preprocessed. The role matching degree matrix is derived from the similarity calculation of the previous action label and the role action template, the matrix row represents the target individual, the column represents the role label, and the matrix element represents the matching score of the individual and each role. The role priority matrix is derived from the event-driven role priority configuration, and each column in the matrix corresponds to the priority weight of a role label. The system concatenates the two matrices in column vector mode to construct a unified state representation tensor for the input of the deep reinforcement learning strategy network. The specific concatenation method is to first normalize the role matching degree matrix by column to compress all scores to the interval of 0 to 1, to ensure the numerical stability of the input and the convergence of the model; then the normalized role priority matrix is directly concatenated to the right of the role matching degree matrix to form a state tensor, with each target individual corresponding to a row and the column vector length being twice the number of roles. The concatenation order strictly maintains consistency with the order of the role labels to prevent the label correspondence relationship from being disordered. The state tensor is input to the deep reinforcement learning model in batch form.
[0082] After the completion of the state tensor construction, the recognition success rate index in the current face recognition cycle is extracted. The index is output by the aforementioned cycle statistics module, and the specific calculation method is to divide the number of successful recognition times by the total number of recognition attempts in the cycle, and keep two decimal places of accuracy (for example, 0.75 represents 75%). The system then counts the proportion of each role label in the recognition success samples according to the "role label" field in the target individual structure. For example, if the teacher role appears 6 times, the student role appears 3 times, and the parent role appears 1 time in 10 recognition success samples, the role composition ratio is 60%, 30%, and 10%, respectively. In order to combine the recognition success rate index with the composition ratio of the role label, the system takes the appearance ratio of the role label as a weight factor, multiplies it by the cycle recognition success rate index, and obtains the reward value of each role. For example, if the recognition success rate index is 0.75, the reward value of the teacher is 0.75x0.6=0.45, the reward value of the student is 0.75x0.3=0.225, and the reward value of the parent is 0.75x0.1=0.075. All reward values are kept to two decimal places in vector form and stored uniformly in the role reward vector as the reward input of the subsequent reinforcement learning model.
[0083] Based on the constructed state tensor and role reward vector, the system introduces a policy gradient method to construct a deep reinforcement learning model for dynamic optimization of role priority weights. The model structure used is a multi-layer perception network, which includes an input layer, at least two hidden layers, and an output layer. The number of nodes in the input layer is equal to the column vector dimension of the state tensor, and the number of nodes in the output layer is equal to the number of role labels. The model connects each layer through a nonlinear activation function (such as ReLU) to ensure the fitting ability of the model. The policy function outputs the priority selection probability of each role in the form of a probability distribution, and uses a softmax layer to normalize the output to avoid numerical overflow and excessive bias. The reward signal is input into the loss function to maximize the expected reward as the target function of policy gradient optimization. The system outputs a probability distribution vector of role priority configuration according to the current state input, which will be converted into actual priority weights in the subsequent steps and used to update the original role priority matrix.
[0084] A gradient descent-based network parameter optimization process is performed on the constructed deep reinforcement learning model. First, the batch state input and the role reward value input are input into the model, the policy output and the corresponding loss function are calculated, the loss function adopts the policy gradient loss plus the L2 regularization term, which ensures the generalization ability of the model. The system performs back propagation calculation on each layer weight matrix and bias vector of the model according to the loss function, and updates the network parameters layer by layer. In order to prevent gradient explosion or gradient disappearance, the system sets the gradient clipping threshold to 1.0, and the gradient exceeding the value will be truncated in the update. After each round of training, the system judges the convergence of the model through the loss value drop curve, and when the loss value change is less than 0.001 for three periods, it is determined that the model converges. After convergence, the system extracts the role priority weight from the policy output, converts the probability distribution vector into a new priority weight according to the role order, replaces the current role priority matrix, and completes the automatic optimization update process of the role priority configuration as the input of the next identification period.
[0085] Embodiment 2, the difference between embodiment 2 and embodiment 1 of the application is that the embodiment is to introduce the community personnel multi-modal recognition system based on the image perception large model.
[0086] Figure 2 The structural diagram of the community personnel multi-modal recognition system based on the image perception large model is given, the community personnel multi-modal recognition system based on the image perception large model includes a role priority construction module, a target extraction module, an action labeling module, a role assignment module and a role priority adjustment module, wherein:
[0087] The role priority construction module: according to the timestamp and shooting position of the image frame, the community event scheduling log is retrieved to determine the event label, and the identification target role list and priority weight are set according to the event label, and the role priority matrix corresponding to the current image frame is constructed;
[0088] The target extraction module: according to the community monitoring image frame sequence, the size standardization processing is carried out, the boundary box and the mask information of all detectable target individuals are extracted by inputting the deep target recognition network model, and the multi-target image entity set is constructed across frames continuously;
[0089] The action labeling module: the image sequence of the individuals in the target image entity set is cut and the optical flow tensor is calculated between frames, the short-time behavior feature is constructed combined with the boundary box track, and the action class label is generated by inputting into the time sequence action recognition network;
[0090] The role assignment module: based on the event role list, the action template is constructed, the semantic embedding and template matching of the target individual action label are carried out, the role matching degree matrix is generated, and the face recognition order of the target is sorted according to the product of the matching degree and the role priority matrix;
[0091] The role priority adjustment module: a recognition success rate of a recognition period is counted, a role matching degree matrix and a role priority matrix are combined to construct a state input tensor, a reward signal is generated in combination with the recognition success rate, a strategy gradient type reinforcement learning network is trained, and a role priority weight adjustment result is output to update the role priority matrix.
[0092] The above formulas are dimensionless values for calculation, the formulas are obtained by software simulation of a large amount of data to obtain a formula of the latest real situation, and the preset parameters and threshold values in the formula are set by a person skilled in the art according to the actual situation.
[0093] The above embodiments can be realized wholly or partially by software, hardware, firmware or any other combination. When realized by software, the above embodiments can be realized wholly or partially in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are wholly or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable devices. The computer instructions can be stored in a computer readable storage medium or transferred from one computer readable storage medium to another, for example, the computer instructions can be transferred from one website, computer, server or data center to another by wired (for example, infrared, wireless, microwave, etc.) mode. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center and the like containing one or more available medium sets. The available medium can be a magnetic medium (for example, floppy disk, hard disk, magnetic tape), an optical medium (for example, DVD) or a semiconductor medium. The semiconductor medium can be a solid state disk.
[0094] Those skilled in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0095] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working process of the above-described system, device and module can refer to the corresponding process in the foregoing method embodiments, which will not be described here.
[0096] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the division of the above-described device embodiments is merely a logical function division, and there can be another division manner for the actual implementation, for example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different modules can be indirect couplings or communication connections through some interfaces, devices or modules, and can be electrical, mechanical or other forms.
[0097] The modules illustrated as separated components can or can not be physically separated, and the components illustrated as modules can or can not be physical modules, and can be located in one place, or can be distributed on multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments.
[0098] In addition, the function modules in each embodiment of the present application can be integrated into a processing module, or each module can be physically present alone, or two or more modules can be integrated into one module.
[0099] If the functions are realized in the form of software function modules and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the parts of the prior art that make contributions or parts of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program code storage media.
[0100] The above is merely specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0101] Finally: the above only for the preferred embodiments of the present application, and not for limiting the present application, any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application, should be included in the scope of protection of the present application.
Claims
1. A community personnel multi-modal recognition method based on an image perception large model, characterized in that, Comprise the following steps: S1: parse the community event label of the image frame from the community monitoring image, preset event-driven role priority matrix; S2: extract all detectable target individuals in the continuous frames of the monitoring image, perform bounding box positioning and instance segmentation on the target individuals, and construct a multi-target image entity set; S3: input the target individual structure in the target image entity set into the time sequence action recognition network, extract the behavior segment of the target individual based on the optical flow feature and dynamic trajectory, and perform action category labeling; S4: establish the action template label of different roles, obtain the role matching degree matrix according to the semantic matching of the target individual action label and the action template label, and perform role assignment and face recognition order calibration on the target individual based on the role priority matrix and the role matching degree matrix; S5: based on the face recognition order calibration result, execute face recognition scheduling in the face recognition cycle, and calculate the face recognition success rate; S6: based on the face recognition success rate, obtain the role recognition success rate, take the role priority matrix and the role matching degree matrix as the state space, and take the role recognition success rate as the reward to construct a reinforcement learning strategy network, update the weight value of the role priority matrix as the input of the next face recognition cycle.
2. The community personnel multi-modal identification method based on image perception large model according to claim 1, characterized in that, In S1, the community event label of the image frame is parsed from the community monitoring image, and the event-driven role priority matrix is preset, which specifically comprises: Extract the timestamp and image shooting position of the image frame from the community monitoring image as the event binding condition; According to the regional nature of the community, preset the event type label, retrieve the community event scheduling log based on the event binding condition, and identify the event type label to which the current image belongs; According to the event type label, set the identification target role list of the event and the corresponding identification priority, and generate the role priority weight value set under the context of the current image frame; Based on the association structure of the role list and the role priority weight value set, construct the role priority matrix corresponding to different event labels.
3. The community personnel multi-modal identification method based on image perception large model according to claim 1, characterized in that, In S2, all detectable target individuals in the continuous frames of the monitoring image are extracted, the bounding box positioning and instance segmentation are performed on the target individuals, and the multi-target image entity set is constructed, which specifically comprises: Perform pixel matrix size alignment on the community monitoring image frame sequence, and convert all image frame data to a standardized single color channel format; Input the standardized image frame sequence into the deep target recognition network model, output the bounding box of all target individuals, perform mask generation on the bounding box area, and extract the pixel mask graph matched with the target individual contour; Integrate the corresponding target individual bounding box area, mask graph data and time index in the standardized image frame into the target individual structure; Organize all target individual structures in chronological order to form a cross-frame continuous multi-target image entity set.
4. The community personnel multi-modal identification method based on image perception large model according to claim 1, characterized in that, In S3, the target individual structure in the target image entity set is input into the time sequence action recognition network, the behavior segment of the target individual is extracted based on the optical flow feature and dynamic trajectory, and the action category labeling is performed, which specifically comprises: Perform interframe cutting on the bounding box area of each target individual in the multi-target image entity set, and construct the image frame sequence of the corresponding target individual; Perform inter-frame optical flow calculation on the target individual image frame sequence, and generate an optical flow tensor feature composed of pixel-level motion vectors; Splice the target individual image frame sequence into a bounding box track, and jointly construct a short-time behavior feature of the target individual based on the optical flow tensor feature and the bounding box track; Input the short-time behavior feature into the time sequence action recognition network, output the action class probability distribution, and write the action label with the maximum probability back to the target individual structure corresponding to the multi-target image entity set.
5. The community personnel multi-modal identification method based on image perception large model according to claim 1, characterized in that, In S4, action template labels of different roles are established, and a role matching degree matrix is obtained according to the semantic matching of the target individual action label and the action template label. The role assignment and face recognition order calibration of the target individual based on the role priority matrix and the role matching degree matrix specifically includes: Obtain the role list corresponding to the current event label, and construct the role action template; Perform semantic embedding processing on the action label of the target individual and the role action template, and generate the action vector and the role action template vector matched in the same semantic space; Calculate the vector similarity of the action vector and the role action template vector in the semantic space, and output the role matching degree matrix according to the calculation result; Take the role label corresponding to the maximum matching degree as the role assignment of the target individual in the current event context, and write it into the corresponding target individual structure; Based on the matrix element product of the role matching degree matrix and the role priority matrix, the face recognition priority order of the target individual is calibrated in descending order of the result.
6. The community personnel multi-modal identification method based on image perception large model according to claim 1, characterized in that, In S5, face recognition scheduling is performed in the face recognition period based on the face recognition order calibration result, and the face recognition success rate is calculated, specifically including: Based on the real-time acquired community monitoring image, the face recognition period with variable length is divided according to the number of detectable target individuals in the image, and the target individuals in the monitoring image are scheduled for face recognition according to the face recognition order calibration; Wherein, each face recognition period takes the first successful recognition as the end marker, and the face recognition success rate is calculated based on the total number of face recognition attempts and the number of successful attempts in each face recognition period.
7. The community personnel multi-modal identification method based on image perception large model according to claim 1, characterized in that, In S6, a reinforcement learning strategy network is constructed, and the weight value of the role priority matrix is updated, specifically including: Collect the role matching degree matrix and the role priority matrix generated in the current face recognition period, splice and construct a state representation tensor in the form of column vectors as the state input of the strategy network; Obtain the face recognition success rate index in the current face recognition period, obtain the role label with successful recognition based on the proportion of the role label, and take the recognition success rate index as the reward value of the corresponding role label; Establish a deep reinforcement learning model based on the change of strategy gradient, take the state tensor as the strategy input and the role label reward value as the reward, and construct an update strategy function of the role priority weight; Perform network parameter optimization process based on gradient descent, and perform round-by-round training and back propagation update of the weight matrix and bias vector of the multi-layer perception structure of the strategy function; After the strategy network converges, output the role priority weight update result under the current state, and replace the original role priority matrix.
8. The community staff multi-modal recognition system based on image perception large model, used to implement the community staff multi-modal recognition method based on image perception large model according to any one of claims 1-7, characterized in that, The system includes a role priority construction module, a target extraction module, an action labeling module, a role assignment module, and a role priority adjustment module, wherein: A role priority construction module: according to the timestamp and shooting position of the image frame, the community event scheduling log is retrieved to determine the event label, and the identification target role list and priority weight are set according to the event label, and a role priority matrix corresponding to the current image frame is constructed; A target extraction module: according to the community monitoring image frame sequence, size standardization processing is performed, the deep target recognition network model is input to extract the boundary box and mask information of all detectable target individuals, and a cross-frame continuous multi-target image entity set is constructed; An action labeling module: the image sequence of the individuals in the target image entity set is cut and the optical flow tensor is calculated between frames, the short-time behavior characteristics are constructed combined with the boundary box track, and the action class label is generated by inputting into the time sequence action recognition network; A role assignment module: based on the event role list, an action template is constructed, the target individual action label is semantically embedded and template matched, a role matching degree matrix is generated, and the target is sequentially sorted by face recognition according to the product of the matching degree and the role priority matrix; A role priority adjustment module: the recognition success rate of the recognition cycle is counted, the state input tensor is constructed with the role matching degree matrix and the role priority matrix, the reward signal is generated combined with the recognition success rate, the strategy gradient type reinforcement learning network is trained, and the role priority weight adjustment result is output to update the role priority matrix.
Citation Information
Patent Citations
Crowd behavior recognition method based on social adaptation model
CN112329539A
Community governance system combining image analysis and big data deep learning
CN118917980A