A large scene character behavior analysis method, system, device and storage medium
By combining multiple long-range cameras with time synchronization and multimodal analysis, the problem of single-camera data being easily affected by viewpoint obstruction in existing technologies is solved. Through multimodal analysis, the accuracy and flexibility of human behavior analysis in large-scale scenes in existing technologies are solved, achieving efficient and accurate object attention analysis.
Patent Information
- Application Number
- CN202511031897.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2045-07-25
AI Technical Summary
When analyzing human behavior in large-scale scenes, existing technologies are prone to the influence of single-camera data due to the obstruction of the viewpoint. They cannot accurately capture attention analysis of small or densely placed objects, resulting in poor accuracy. They also cannot capture subtle behavioral clues such as the direction of gaze or body orientation, leading to low analysis accuracy.
Multiple long-range cameras are used for scene area calibration. Combined with time synchronization and multimodal analysis, a pedestrian re-identification model is used to track the trajectory of people. An adaptive threshold adjustment mechanism and semantic segmentation technology are introduced to dynamically adapt to scene changes and integrate multimodal data for analysis.
It significantly improves the accuracy and robustness of attention analysis for small items, enhances the accuracy and flexibility of human behavior analysis in large-scale scenarios, and reduces maintenance costs and system maintenance costs.
Smart Images

Figure CN120525971B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence, and in particular to a large scene character behavior analysis method, system, device and storage medium. BACKGROUND
[0002] With the increasing degree of social informatization and digitization, the demand for character behavior analysis in large scenes (such as public places, transportation hubs, industrial production sites, etc.) is growing. In the prior art, character behavior is generally analyzed based on character trajectories, such as recording character movement trajectories and dwell time in large scenes (such as furniture city, automobile 4S store) through monitoring cameras to infer attention. Such systems use simple motion detection or object tracking algorithms, such as the passenger flow analysis system of Hikvision or Dahua.
[0003] However, the prior art relies mainly on single camera data, which is easily affected by visual angle occlusion, so the attention analysis effect is poor for small or densely placed objects (such as cosmetics, electronic products), the accuracy is low, and subtle behavior clues such as gaze direction or body orientation cannot be captured, resulting in poor accuracy of character behavior analysis in large scenes. SUMMARY
[0004] In order to improve the accuracy of character behavior analysis in large scenes, the present application provides a large scene character behavior analysis method, system, device and storage medium.
[0005] In a first aspect, the present application provides a large scene character behavior analysis method, which adopts the following technical solution:
[0006] The scene area of the plurality of long-range cameras is calibrated and a plurality of area calibration information is recorded, the area calibration information including the coverage area of the plurality of long-range cameras and the object information and object coordinates in the coverage area;
[0007] A plurality of scene data is collected by the plurality of long-range cameras, and the plurality of scene data is time-synchronized;
[0008] The pedestrian re-identification model is used to track the character trajectory in the plurality of scene data to obtain a plurality of area dwell times and a plurality of area images corresponding to the plurality of area dwell times;
[0009] A plurality of character analysis results is obtained by performing multi-modal analysis on the plurality of area images;
[0010] The plurality of character analysis results is fused to obtain a comprehensive character analysis result.
[0011] By the technical solution, the character behavior analysis is performed on the characters by using the multi-path long-range cameras in a large scene, time synchronization is combined, multi-source data integration is optimized, the accuracy of small item attention analysis is significantly improved, and the analysis robustness is improved by multi-camera fusion.
[0012] In one specific implementation, the method is further based on a semantic segmentation model, and after the scene area calibration of the plurality of long-range cameras and the recording of the plurality of area calibration information, the method further includes:
[0013] The semantic segmentation model is used to identify the scene item boundary;
[0014] It is determined whether the item boundary changes;
[0015] If the item boundary changes, the area calibration of the plurality of long-range cameras is updated, and the plurality of updated area calibration information is recorded.
[0016] By the technical solution, the automatic calibration is realized, and the maintenance cost is reduced. Combined with the semantic segmentation technology, the system flexibility and accuracy are improved by dynamically adapting to the shelf movement or item adjustment. All dynamic area calibration methods based on computer vision (such as semantic segmentation and target detection) and their applications in behavior analysis are covered.
[0017] In one specific implementation, the tracking of the character trajectory in the plurality of scene data by using the pedestrian re-identification model to obtain the plurality of area stay times and the plurality of area images corresponding to the plurality of area stay times includes:
[0018] An adaptive threshold adjustment mechanism is introduced to adjust the stay time threshold according to the scene data;
[0019] The pedestrian re-identification model is used to track the character trajectory in the plurality of scene data to obtain the plurality of area stay times;
[0020] It is determined whether the plurality of area stay times exceeds the stay time threshold;
[0021] If the plurality of area stay times exceeds the stay time threshold, the plurality of area stay times and the plurality of area images corresponding to the plurality of area stay times are recorded.
[0022] In one specific implementation, the adaptive threshold adjustment mechanism is introduced to adjust the stay time threshold according to the scene includes:
[0023] The plurality of scene data is analyzed to obtain passenger flow data;
[0024] The stay time threshold is adjusted according to the passenger flow data, and the calculation formula is:
[0025] ;
[0026] wherein T is a threshold of dwell time, T0 is a base threshold, D is a crowd density, and a is an adjustment coefficient.
[0027] Through the above technical solution, the application improves the robustness and accuracy of behavior analysis by dynamically adjusting the threshold. The crowd density is introduced as a dynamic variable, and the automatic adjustment is realized by combining mathematical formulas, which is better than the existing static threshold method. All algorithms and applications for dynamically adjusting the dwell time threshold based on crowd density or other scene parameters (such as light and time period) are covered.
[0028] In one specific implementation, the method is also based on a pose estimation model, a lightweight model, and MediaPipe, and the multi-modal analysis of the plurality of region images to obtain a plurality of person analysis results includes:
[0029] using the pose estimation model to analyze the plurality of region images to obtain a plurality of person face orientation confidence corresponding to the plurality of region images;
[0030] using the lightweight model to analyze the plurality of region images to obtain a plurality of gaze confidence corresponding to the plurality of region images;
[0031] using the MediaPipe to analyze the plurality of region images to obtain a plurality of hand action confidence corresponding to the plurality of region images;
[0032] performing multi-modal analysis according to the plurality of person face orientation confidence, the plurality of gaze confidence, the plurality of hand action confidence, and the plurality of region calibration information to obtain a plurality of person analysis results.
[0033] In one specific implementation, the multi-modal analysis according to the plurality of person face orientation confidence, the plurality of gaze confidence, the plurality of hand action confidence, and the plurality of region calibration information to obtain a plurality of person analysis results includes:
[0034] receiving a plurality of preset weight information, the weight information including face orientation weight, gaze weight, and hand action weight;
[0035] performing multi-modal analysis according to the plurality of weight information, the plurality of person face orientation confidence, the plurality of gaze confidence, the plurality of hand action confidence, and the plurality of region calibration information to obtain a plurality of person analysis results, the person analysis results including single camera attention confidence, and the calculation formula is:
[0036] ;
[0037] wherein C is a single camera confidence of the person paying attention to the object; is a face orientation confidence; is a gaze confidence; is a hand action confidence; W1, W2, and W3 are a face orientation weight, a gaze weight, and a hand action weight, respectively.
[0038] Through the above technical solution, the application fuses multi-modal data to improve the accuracy of small object attention analysis (up to 90%). The three behavior clues are integrated and the weight is dynamically adjusted to adapt to different scenes and object types, which is better than single modal analysis.
[0039] In a specific implementable embodiment, the fusing of the plurality of person analysis results to obtain a person analysis comprehensive result includes:
[0040] The fusing of the plurality of person analysis results to obtain a person analysis comprehensive result includes a final comprehensive confidence, and the calculation formula is:
[0041]
[0042] wherein, is a final comprehensive confidence, indicating the attention probability of the person to a specific object or area; Ci is a single camera confidence of the i-th camera; Wi is a weight of the i-th camera, reflecting the reliability or importance of the camera data; N is the number of cameras participating in fusion; is the sum of all camera weights.
[0043] Through the above technical solution, the application improves the analysis robustness through multi-camera fusion. Combined with time synchronization and view weight, the multi-source data integration is optimized, and the accuracy of small object attention analysis is significantly improved. All multi-camera data fusion-based behavior analysis methods are covered, including time synchronization and confidence weighting mechanism.
[0044] In a second aspect, the application provides a large scene person behavior analysis system, which adopts the following technical solution: the system includes:
[0045] A scene area calibration module is configured to calibrate a scene area of the plurality of multi-channel long-range cameras and record a plurality of area calibration information, wherein the area calibration information includes an object information and an object coordinate in a coverage area of the multi-channel long-range cameras.
[0046] A scene data acquisition module is configured to acquire a plurality of scene data by using the plurality of multi-channel long-range cameras, and time-synchronize the plurality of scene data.
[0047] a person trajectory tracking module, configured to track person trajectories in the plurality of scene data by using a pedestrian re-identification model to obtain a plurality of region stay times and a plurality of region images corresponding to the plurality of region stay times;
[0048] a person behavior analysis module, configured to perform multi-modal analysis on the plurality of region images to obtain a plurality of person analysis results;
[0049] a person behavior analysis fusion module, configured to fuse the plurality of person analysis results to obtain a person analysis comprehensive result.
[0050] In a third aspect, the present application provides a computer device, which adopts the following technical scheme: comprising a memory and a processor, the memory stores a computer program capable of being loaded and executed by the processor to perform the above-mentioned large-scene person behavior analysis method.
[0051] In a fourth aspect, the present application provides a computer readable storage medium, which adopts the following technical scheme: storing a computer program capable of being loaded and executed by the processor to perform the above-mentioned large-scene person behavior analysis method.
[0052] In summary, the present application has the following beneficial technical effects:
[0053] (1) The present application proposes a complete behavior analysis system integrating pedestrian re-identification, pose estimation, multi-modal analysis and dynamic calibration, realizes multi-technology integration, forms a complete behavior analysis chain, uses multiple long-range cameras to analyze the behavior of persons in a large scene, combines time synchronization, optimizes multi-source data integration, significantly improves the accuracy of small item attention analysis, and improves the analysis robustness through multi-camera fusion
[0054] (2) The present application realizes automatic calibration and reduces maintenance cost. Combined with the semantic segmentation technology, the system flexibility and accuracy are improved, and the accuracy of person behavior analysis in a large scene is improved.
[0055] (3) The present application improves the robustness and accuracy of behavior analysis by dynamically adjusting the threshold. The crowd density is introduced as a dynamic variable, and the automatic adjustment is realized by combining mathematical formula, which is superior to the existing static threshold method. All algorithms and applications based on crowd density or other scene parameters (such as light and time period) for dynamically adjusting stay time threshold are covered, which improves the accuracy of person behavior analysis in a large scene.
[0056] (4) The present application fuses multi-modal data to improve the precision of small item attention analysis (up to 90%). Three kinds of behavior clues are integrated and the weight is dynamically adjusted, which is suitable for different scenes and item types, is superior to single modal analysis, and improves the accuracy of person behavior analysis in a large scene.
[0057] (5) Through the automation process and low cost implementation (using existing monitoring system), improve the efficiency and data value of the business scenario. BRIEF DESCRIPTION OF DRAWINGS
[0058] Fig. 1 is a flow chart of a large scene character behavior analysis method in an embodiment of the present application.
[0059] Fig. 2 is a structural block diagram of a large scene character behavior analysis method in an embodiment of the present application.
[0060] Reference signs: 201, scene area calibration module; 202, scene data acquisition module; 203, character trajectory tracking module; 204, character behavior analysis module; 205, character behavior analysis fusion module. DETAILED DESCRIPTION
[0061] The following will be combined with the accompanying Figs. 1-2 Further detailed description of the present application.
[0062] Embodiments of the present application disclose a large scene character behavior analysis method, which is used to improve the accuracy of character behavior analysis in a large scene.
[0063] With the improvement of social informatization and digitization, the demand for character behavior analysis in a large scene (such as public places, transportation hubs, industrial production sites, etc.) is increasing. In the prior art, character behavior is generally analyzed based on character trajectory, such as recording character moving trajectory and staying time through a monitoring camera in a large scene (such as a furniture city, a car 4S store) to infer attention. Such a system uses simple motion detection or object tracking algorithm, for example, the passenger flow analysis system of Hikvision or Dahua.
[0064] However, the prior art relies on single camera data, which is easily affected by visual angle obstruction, so the attention analysis effect is poor for small or densely placed objects (such as cosmetics, electronic products), the accuracy is low, and subtle behavior clues such as gaze direction or body orientation cannot be captured. In addition, it has strong limitations, relies on fixed threshold staying time judgment, lacks adaptability to scene dynamic changes, is easily affected by crowd density or light changes, and leads to poor accuracy of character behavior analysis in a large scene.
[0065] Therefore, the present application proposes a large scene character behavior analysis method, which is used to improve the accuracy of character behavior analysis in a large scene.
[0066] The hardware part of the present application includes:
[0067] Multi-channel long-range camera: covering different areas in the scene, pre-calibrating object coordinates, supporting dynamic area adjustment.
[0068] Processing server: equipped with high-performance GPU or other inference devices, running face recognition, pedestrian re-identification, pose estimation, and multi-modal attention analysis model.
[0069] Background database: stores item information, attention records, and behavior analysis results.
[0070] As shown in Fig. 1 The method comprises the following steps:
[0071] S10, calibrate the scene area of a plurality of multi-lens long-range cameras and record a plurality of area calibration information, the area calibration information including the coverage area of the multi-lens long-range camera and the item information and item coordinates in the coverage area.
[0072] Specifically, a plurality of multi-lens long-range cameras are arranged in a large scene, each multi-lens long-range camera covers part of the area in the scene, and each area has corresponding items. For example, in a retail store scene, each multi-lens long-range camera covers an area containing shelves and goods. First, calibrate the scene area of a plurality of multi-lens long-range cameras and record a plurality of area calibration information, the area calibration information including the coverage area of the multi-lens long-range camera and the item information and item coordinates in the coverage area.
[0073] S20, collect a plurality of scene data using a plurality of multi-lens long-range cameras, and time synchronize the plurality of scene data.
[0074] Specifically, each multi-lens long-range camera collects scene data of the corresponding area. Before analyzing the collected plurality of scene data, the plurality of scene data is time synchronized.
[0075] S30, track the trajectory of a person in the plurality of scene data using a pedestrian re-identification model to obtain a plurality of area stay times and a plurality of area images corresponding to the plurality of area stay times.
[0076] Specifically, the collected plurality of scene data is processed, a stay time threshold is preset, and a pedestrian re-identification model (such as DeepSORT or FairMOT) is used to track the trajectory of a person according to the stay time threshold to obtain a plurality of area stay times and a plurality of area images corresponding to the plurality of area stay times.
[0077] S40, perform multi-modal analysis on the plurality of area images to obtain a plurality of person analysis results.
[0078] Specifically, multi-modal attention analysis is introduced to combine face orientation, gaze estimation (based on eye key points), and hand movement (detecting hand pointing or touching action) to perform multi-modal analysis on the plurality of area images to obtain a plurality of person analysis results.
[0079] S50, fusing the results of the character analysis to obtain a comprehensive character analysis result.
[0080] Specifically, the coverage areas of different multi-lane long-range cameras overlap, and the same dwell time may be arranged by different multi-lane long-range cameras to obtain images of areas with different perspectives, so the results of the character analysis of the multi-lane cameras in the same period are fused to obtain a comprehensive character analysis result.
[0081] In large scenes, the present application uses multi-lane long-range cameras to analyze the behavior of characters, combines time synchronization, optimizes multi-source data integration, significantly improves the accuracy of small item attention analysis, and improves the analysis robustness through multi-camera fusion.
[0082] In one embodiment, to improve the accuracy of character behavior analysis in large scenes, after calibrating the scene area of several multi-lane long-range cameras and recording several area calibration information, the following steps can be performed:
[0083] First, the semantic segmentation model is used to identify the scene item boundary, specifically, the semantic segmentation model is used to identify the scene item boundary, for example, in a retail store scene, a semantic segmentation model (such as DeepLabv3) is used to identify the boundaries of shelves and items in real time;
[0084] Then, determine whether the scene item boundary has changed, if the scene item boundary has changed, update the area calibration of the several multi-lane long-range cameras, and record the updated several area calibration information, specifically, determine whether the item boundary has changed, for example, in a retail store scene, the shelves move or the buyer takes the goods on the shelves, the scene item boundary will change; To improve the accuracy of character behavior analysis in large scenes, dynamic area calibration is introduced, that is, when the scene item boundary changes are detected, the multi-lane long-range camera area calibration information is updated regularly, for example, in a retail store scene, when the real-time identification of shelf and item boundary changes, update the area coordinates every hour, the error is controlled within ±5cm, automatically generate the mapping relationship between the area and the item, and store it in the database.
[0085] The present application realizes automatic calibration and reduces maintenance cost. Combined with semantic segmentation technology, dynamically adapt to shelf movement or item adjustment, improve system flexibility and accuracy. Cover all dynamic area calibration methods based on computer vision (such as semantic segmentation, target detection) and their applications in behavior analysis.
[0086] In one embodiment, to improve the accuracy of character behavior analysis in large scenes, the step of tracking the character trajectory in the several scene data using the pedestrian re-identification model to obtain the several area dwell time and the several area images corresponding to the several area dwell time can be specifically performed as:
[0087] Firstly, the adaptive threshold adjustment mechanism is introduced to adjust the stay time threshold according to the scene data. Specifically, the crowd density is introduced as a dynamic variable, and the stay time threshold is adjusted adaptively according to the scene data.
[0088] Then, the pedestrian re-identification model is used to track the trajectories of the persons in the scene data to obtain the stay time of each region. Specifically, the pedestrian re-identification model (such as DeepSORT or FairMOT) is used to track the trajectories of the persons in the scene data to obtain the stay time of each region.
[0089] Next, it is determined whether the stay time of each region exceeds the stay time threshold. If the stay time of each region exceeds the stay time threshold, the stay time of each region and the image corresponding to the stay time of each region are recorded. Specifically, during the collection process, each stay time of each region collected by each multi-view long-range camera needs to be compared with the stay time threshold. Only when the stay time of each region exceeds the stay time threshold, the stay time of each region and the image corresponding to the stay time of each region are recorded. After comparing all the stay times of all the regions collected by all the multi-view long-range cameras, the stay time of each region and the image corresponding to the stay time of each region are obtained.
[0090] In one embodiment, in order to improve the accuracy of the analysis of the behavior of the persons in a large scene, the adaptive threshold adjustment mechanism is introduced to adjust the stay time threshold according to the scene data. This step can be specifically implemented as follows:
[0091] Firstly, the traffic data is obtained by analyzing the scene data. Specifically, in reality, when the peak period of the crowd flow occurs, the stay time of the persons will be longer. Therefore, the traffic data is obtained by analyzing the collected scene data.
[0092] Then, the stay time threshold is adjusted according to the traffic data. The calculation formula is as follows:
[0093] ;
[0094] Wherein, T is the stay time threshold, T0 is the basic threshold, D is the crowd density, and a is the adjustment coefficient.
[0095] Specifically, the stay time threshold is dynamically adjusted based on the real-time traffic statistics. The calculation formula of the stay time threshold is as follows:
[0096] ;
[0097] Wherein, T0 is the basic threshold (such as 5 seconds), D is the crowd density (normalized value 0-1), and a is the adjustment coefficient (0.5).
[0098] T: Actual usage of the threshold of the dwell time (unit: seconds) to determine whether the person stays in a certain area for a long enough time to trigger subsequent behavior analysis (such as pose estimation).
[0099] T0: The basic threshold of the dwell time, set to 5 seconds, representing the default threshold in a standard scenario (without special crowd density influence).
[0100] D: Crowd density, a normalized value between 0 and 1, representing the degree of crowd concentration in the scene.
[0101] D=0: The scene is almost empty (such as a retail store at night).
[0102] D=1: The scene is very crowded (such as the peak of a promotional event).
[0103] The crowd density can be detected by the camera in real time and normalized (for example, the number of people / maximum capacity).
[0104] α: Adjustment coefficient, set to 0.5, used to control the degree of influence of crowd density on the threshold. The smaller the value, the smaller the adjustment range of the threshold; the larger the value, the larger the adjustment range.
[0105] The present application improves the robustness and accuracy of behavior analysis by dynamically adjusting the threshold. By introducing crowd density as a dynamic variable and combining mathematical formulas to achieve automatic adjustment, it is superior to existing static threshold methods. It covers all algorithms and applications that dynamically adjust the dwell time threshold based on crowd density or other scene parameters (such as light, time period).
[0106] In one embodiment, in order to improve the accuracy of the behavior analysis of the characters in a large scene, the step of performing multi-modal analysis on the plurality of region images to obtain a plurality of character analysis results can be specifically implemented as:
[0107] First, the pose estimation model is used to analyze the plurality of region images to obtain a plurality of face orientation confidence corresponding to the plurality of region images. Specifically, for the region image exceeding the threshold, the pose estimation model (such as MoveNet) is input to detect the human body key points (left and right shoulders, ears, eyes, a total of 6 points), and the face orientation angle (error ± 5°) is calculated from the human body key points. The face orientation angle is converted into a face orientation confidence, reflecting whether the face of the character is facing the target object.
[0108] Second, the lightweight model is used to analyze the region image to obtain a plurality of gaze confidence corresponding to the plurality of region images. Specifically, the eye key points are analyzed by the lightweight model (such as GazeNet), and the gaze estimation is obtained based on the eye key points analysis, and the gaze estimation is converted into a gaze confidence, reflecting whether the gaze is pointing to the target object.
[0109] Then, the region image is analyzed by using the MediaPipe to obtain a plurality of hand action confidences corresponding to the plurality of region images. Specifically, the hand action is analyzed by using the MediaPipe to detect the region image, whether the hand points or touches the target object is analyzed, the hand action is converted into the hand action confidence, and whether the hand points or touches the target object is reflected.
[0110] Then, multi-modal analysis is performed according to the plurality of person face orientation confidences, the plurality of gaze confidences, the plurality of hand action confidences, and the plurality of region calibration information to obtain a plurality of person analysis results. Specifically, after the face orientation angle, the gaze estimation, and the hand action of the person are obtained, the region calibration information of the multi-lens long shot camera (which includes the information of the objects in the region) needs to be combined for analysis to obtain the person analysis result, that is, the comprehensive confidence of the single camera paying attention to the object of the person. After all the multi-lens long shot cameras are analyzed, a plurality of person analysis results are obtained.
[0111] In one embodiment, in order to improve the accuracy of the person behavior analysis in a large scene, the step of performing multi-modal analysis according to the plurality of person face orientation confidences, the plurality of gaze confidences, the plurality of hand action confidences, and the plurality of region calibration information to obtain a plurality of person analysis results can be specifically implemented as:
[0112] First, a plurality of preset weight information is received. The weight information includes a face orientation weight, a gaze weight, and a hand action weight. Specifically, according to the scene requirements (such as different object types or camera angles), the proportion of the face orientation weight, the gaze weight, and the hand action weight of each multi-lens long shot camera is preset.
[0113] Then, multi-modal analysis is performed according to the plurality of weight information, the plurality of person face orientation confidences, the plurality of gaze confidences, the plurality of hand action confidences, and the plurality of region calibration information to obtain a plurality of person analysis results. The person analysis result includes a single camera attention confidence, and the calculation formula is:
[0114] ;
[0115] Wherein, C is the comprehensive confidence of the single camera paying attention to the object of the person; is the face orientation confidence; is the gaze confidence; is the hand action confidence; W1, W2, and W3 are the face orientation weight, the gaze weight, and the hand action weight, respectively.
[0116] Specifically, after obtaining the face orientation angle, gaze estimation and hand action of the person, it is still necessary to combine the regional calibration information of the multi-path long-range camera (which includes the information of the objects in the region) to obtain the person analysis result, i.e. the comprehensive confidence of the single camera focusing on the object, by combining analysis. After analyzing all the multi-path long-range cameras, a plurality of person analysis results are obtained, and the calculation formula of the single camera focus confidence is:
[0117] ;
[0118] Wherein, C: the comprehensive confidence of the single camera focusing on the object, ranging from 0 to 1. The higher the value, the greater the possibility of the person focusing on the object.
[0119] : Face orientation confidence, the face orientation angle calculated based on the human pose estimation model (such as MoveNet), reflecting whether the face of the person is facing the target object, ranging from 0 to 1.
[0120] : Gaze confidence, the gaze estimation model (such as GazeNet) is used to analyze the eye key points to determine whether the gaze is directed to the target object, ranging from 0 to 1.
[0121] : Hand action confidence, the hand action detection model (such as MediaPipe) is used to analyze whether the hand is pointing or touching the target object, ranging from 0 to 1.
[0122] w1, w2, w3: the weights of face orientation, gaze and hand action, respectively, the default values are 0.4, 0.3 and 0.3, respectively, satisfying w1+w2+w3=1, to ensure that the comprehensive confidence C is between 0 and 1.
[0123] Formula effect: The formula fuses the confidence of three modalities (face orientation, gaze and hand action) through weighted summation to comprehensively judge whether the person is focusing on a specific object.
[0124] Multi-modal fusion: A single modality (such as relying only on face orientation) may lead to misjudgment due to angle, light or obstruction, and combining gaze and hand action can improve the robustness and accuracy of the analysis.
[0125] The present application fuses multi-modal data to improve the accuracy of small object focus analysis (up to 90%). The three behavior clues are integrated and the weight is dynamically adjusted to adapt to different scenes and object types, which is better than single modality analysis.
[0126] In one embodiment, in order to improve the accuracy of person behavior analysis in large scenes, the step of fusing a plurality of person analysis results to obtain a person analysis comprehensive result can be specifically implemented as:
[0127] The character analysis results are fused to obtain a character analysis comprehensive result, and the character analysis comprehensive result includes a final comprehensive confidence, and the calculation formula is:
[0128] ;
[0129] Wherein, is the final comprehensive confidence, indicating the attention probability of the character to a certain item or area; Ci is the single camera confidence of the i-th camera; Wi is the weight of the i-th camera, reflecting the reliability or importance of the camera data; N is the number of cameras participating in fusion; is the sum of all camera weights;
[0130] Specifically, the coverage areas of different multi-channel overview cameras will overlap, and the same dwell time may be arranged by different multi-channel overview cameras to obtain area images with different perspectives. Therefore, the character analysis results of the multi-channel cameras in the same period are fused to obtain a character analysis comprehensive result, and the character analysis comprehensive result includes a final comprehensive confidence, and the calculation formula of the final comprehensive confidence is:
[0131] ;
[0132] Wherein, : final comprehensive confidence, indicating the attention probability of the character to a certain item or area, ranging from 0 to 1. The higher the value, the greater the attention possibility.
[0133] Ci: single camera confidence of the i-th camera, calculated by multi-modal attention analysis (i.e. Ci ), reflecting the possibility of the character paying attention to an item in the camera perspective, ranging from 0 to 1.
[0134] Wi: weight of the i-th camera, reflecting the reliability or importance of the camera data. The weight is determined based on the camera perspective quality (such as distance, angle, and blocking condition), and usually ranges from 0 to 1.
[0135] N: number of cameras participating in fusion (for example, 4 overview cameras in a retail scene).
[0136] : sum of all camera weights, used for normalization to ensure between 0 and 1.
[0137] Formula function: this formula fuses the confidence Ci of the multi-channel cameras by weighted average to generate a more accurate and robust comprehensive confidence Cfinal.
[0138] View angle integration: single camera may have inaccurate confidence due to view angle occlusion, light change or distance, fusing multi-camera data can complement the deficiency.
[0139] Weight adjustment: different cameras are given different importance by Wi, for example, cameras with close distance and no occlusion have higher weight.
[0140] Time synchronization: the formula implicitly requires time synchronization to ensure that all Ci correspond to the same period of attention behavior.
[0141] The application improves the robustness of analysis by fusing multi-camera data. Combined with time synchronization and view angle weight, the multi-source data integration is optimized, which significantly improves the accuracy of small item attention analysis. All multi-camera data fusion-based behavior analysis methods are covered, including time synchronization and confidence weighting mechanism.
[0142] Based on the above method, the embodiment of the application also discloses a large scene human behavior analysis system. As Fig. 2 The system includes the following modules:
[0143] The scene area calibration module 201 is used for calibrating the scene area of a plurality of multi-channel long-range cameras and recording a plurality of area calibration information. The area calibration information includes the coverage area of the multi-channel long-range cameras and the item information and item coordinates in the coverage area.
[0144] The scene data acquisition module 202 is used for acquiring a plurality of scene data by using a plurality of multi-channel long-range cameras, and synchronizing the plurality of scene data in time.
[0145] The human trajectory tracking module 203 is used for tracking human trajectories in the plurality of scene data by using a pedestrian re-identification model to obtain a plurality of area stay times and a plurality of area images corresponding to the plurality of area stay times.
[0146] The human behavior analysis module 204 is used for performing multi-modal analysis on the plurality of area images to obtain a plurality of human analysis results.
[0147] The human behavior analysis fusion module 205 is used for fusing the plurality of human analysis results to obtain a human analysis comprehensive result.
[0148] In one embodiment, the human trajectory tracking module 203 is specifically configured to introduce an adaptive threshold adjustment mechanism to adjust the stay time threshold according to the scene data; track human trajectories in the plurality of scene data by using a pedestrian re-identification model to obtain a plurality of area stay times; determine whether the plurality of area stay times exceeds the stay time threshold; if the plurality of area stay times exceeds the stay time threshold, record the plurality of area stay times and a plurality of area images corresponding to the plurality of area stay times.
[0149] In an embodiment, the character trajectory tracking module 203 is specifically configured to parse the scene data to obtain the passenger flow data.
[0150] The stay time threshold is adjusted according to the passenger flow data, and the calculation formula is:
[0151] ;
[0152] Wherein, T is the stay time threshold, T0 is the basic threshold, D is the crowd density, and a is the adjustment coefficient.
[0153] In an embodiment, the character behavior analysis module 204 is specifically configured to analyze the region image by using a pose estimation model to obtain a plurality of character face orientation confidence corresponding to a plurality of region images; analyze the region image by using a lightweight model to obtain a plurality of gaze confidence corresponding to a plurality of region images; analyze the region image by using MediaPipe to obtain a plurality of hand action confidence corresponding to a plurality of region images; and perform multi-modal analysis according to the plurality of character face orientation confidence, the plurality of gaze confidence, the plurality of hand action confidence and a plurality of region calibration information to obtain a plurality of character analysis results.
[0154] In an embodiment, the character behavior analysis module 204 is specifically configured to receive a plurality of preset weight information, the weight information including face orientation weight, gaze weight and hand action weight.
[0155] According to the plurality of weight information, the plurality of character face orientation confidence, the plurality of gaze confidence, the plurality of hand action confidence and the plurality of region calibration information, multi-modal analysis is performed to obtain a plurality of character analysis results, and the character analysis result includes single camera attention confidence, and the calculation formula is:
[0156] ;
[0157] Wherein, C is the comprehensive confidence of the single camera paying attention to the object; is the face orientation confidence; is the gaze confidence; is the hand action confidence; W1, W2 and W3 are the face orientation weight, the gaze weight and the hand action weight respectively.
[0158] In an embodiment, the character behavior analysis fusion module 205 is specifically configured to fuse a plurality of character analysis results to obtain a character analysis comprehensive result, and the character comprehensive analysis result includes a final comprehensive confidence, and the calculation formula is:
[0159] ;
[0160] Wherein, The final overall confidence score represents the probability that a person pays attention to a specific item or area; Ci is the single-camera confidence score of the i-th camera; Wi is the weight of the i-th camera, reflecting the reliability or importance of the data from that camera; N is the number of cameras participating in the fusion. This is the sum of the weights of all cameras.
[0161] This application also discloses a computer device.
[0162] Specifically, the computer device includes a memory and a processor, with the memory storing a computer program that can be loaded and executed by the processor to perform the aforementioned method for analyzing human behavior in a large-scale scene.
[0163] This application also discloses a computer-readable storage medium.
[0164] Specifically, the computer-readable storage medium stores a computer program that can be loaded by a processor and executed, such as the large-scale scene character behavior analysis method described above. The computer-readable storage medium includes, for example, various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0165] This specific embodiment is merely an explanation of the present invention and is not intended to limit the invention. After reading this specification, those skilled in the art can make modifications to this embodiment without contributing any inventive step, but such modifications are protected by patent law as long as they are within the scope of the claims of the present invention.
Claims
1. A method for analyzing behavior of characters in a large scene, characterized by, The method is based on a plurality of multi-lens cameras, a pose estimation model, a lightweight model and MediaPipe, and comprises the following steps: Calibrating a plurality of scene regions of the plurality of multi-lens cameras and recording a plurality of region calibration information, wherein the region calibration information comprises a coverage region of the plurality of multi-lens cameras and object information and object coordinates in the coverage region; Collecting a plurality of scene data by using the plurality of multi-lens cameras, and time-synchronizing the plurality of scene data; Tracking a person trajectory in the plurality of scene data by using a pedestrian re-identification model to obtain a plurality of region stay times and a plurality of region images corresponding to the plurality of region stay times; Performing multi-modal analysis on the plurality of region images to obtain a plurality of person analysis results; Fusing the plurality of person analysis results to obtain a person analysis comprehensive result; The step of tracking a person trajectory in the plurality of scene data by using a pedestrian re-identification model to obtain a plurality of region stay times and a plurality of region images corresponding to the plurality of region stay times comprises the following steps: Introducing an adaptive threshold adjustment mechanism to adjust a stay time threshold according to the plurality of scene data; Tracking a person trajectory in the plurality of scene data by using a pedestrian re-identification model to obtain a plurality of region stay times; Determining whether the plurality of region stay times exceed the stay time threshold; If the plurality of region stay times exceed the stay time threshold, recording the plurality of region stay times and the plurality of region images corresponding to the plurality of region stay times; The step of performing multi-modal analysis on the plurality of region images to obtain a plurality of person analysis results comprises the following steps: Analyzing the plurality of region images by using the pose estimation model to obtain a plurality of person face orientation confidence degrees corresponding to the plurality of region images; Analyzing the plurality of region images by using the lightweight model to obtain a plurality of gaze confidence degrees corresponding to the plurality of region images; Analyzing the plurality of region images by using the MediaPipe to obtain a plurality of hand motion confidence degrees corresponding to the plurality of region images; Performing multi-modal analysis according to the plurality of person face orientation confidence degrees, the plurality of gaze confidence degrees, the plurality of hand motion confidence degrees and the plurality of region calibration information to obtain a plurality of person analysis results; The step of performing multi-modal analysis according to the plurality of person face orientation confidence degrees, the plurality of gaze confidence degrees, the plurality of hand motion confidence degrees and the plurality of region calibration information to obtain a plurality of person analysis results comprises the following steps: Receiving a plurality of preset weight information, wherein the weight information comprises a face orientation weight, a gaze weight and a hand motion weight; Performing multi-modal analysis according to the plurality of weight information, the plurality of person face orientation confidence degrees, the plurality of gaze confidence degrees, the plurality of hand motion confidence degrees and the plurality of region calibration information to obtain a plurality of person analysis results, wherein the person analysis results comprise a single camera attention confidence degree, and the calculation formula is: ; Wherein, C is a single camera on the comprehensive confidence of the character attention to the goods; is a face orientation confidence; is a line of sight confidence; is a hand action confidence; W1, W2, W3 are face orientation weight, line of sight weight and hand action weight, respectively; The step of fusing the plurality of person analysis results to obtain a person analysis comprehensive result comprises the following steps: Fusing the plurality of person analysis results to obtain a person analysis comprehensive result, wherein the person analysis comprehensive result comprises a final comprehensive confidence degree, and the calculation formula is: ; wherein, is the final integrated confidence, representing the probability of the person paying attention to a certain item or area; Ci is the single-camera confidence of the ith camera; Wi is the weight of the ith camera, reflecting the reliability or importance of the data of the camera; N is the number of cameras participating in fusion; is the sum of the weights of all cameras.
2. The method of claim 1, wherein, The method further comprises, based on the semantic segmentation model, after the scene region calibration of the plurality of long-range cameras and the recording of the region calibration information, the following steps: identifying the scene object boundary using the semantic segmentation model; judging whether the scene object boundary changes; if the scene object boundary changes, updating the region calibration of the plurality of long-range cameras and recording the updated region calibration information.
3. The method of claim 2, wherein, The adaptive threshold adjustment mechanism comprises the following steps: analyzing the scene data to obtain the passenger flow data; adjusting the stay time threshold according to the passenger flow data, and the calculation formula is as follows: ; wherein, T is the stay time threshold, T0 is the basic threshold, D is the crowd density, and a is the adjustment coefficient.
4. A large-scale scene character behavior analysis system, characterized in that, The system is based on a plurality of long-range cameras, a pose estimation model, a lightweight model, and MediaPipe, and comprises the following modules: a scene region calibration module (201) for calibrating the scene region of the plurality of long-range cameras and recording the region calibration information, wherein the region calibration information includes the coverage area of the long-range cameras and the object information and object coordinates in the coverage area; a scene data acquisition module (202) for acquiring the scene data using the plurality of long-range cameras and synchronizing the scene data; a person trajectory tracking module (203) for tracking the person trajectory in the scene data using the pedestrian re-identification model to obtain the region stay time and the region image corresponding to the region stay time, and specifically for adjusting the stay time threshold according to the scene data using the adaptive threshold adjustment mechanism; tracking the person trajectory in the scene data using the pedestrian re-identification model to obtain the region stay time; judging whether the region stay time exceeds the stay time threshold; if the region stay time exceeds the stay time threshold, recording the region stay time and the region image corresponding to the region stay time; The person behavior analysis module (204) is configured to perform multi-modal analysis on the region images to obtain person analysis results, specifically configured to analyze the region images by using the pose estimation model to obtain person face orientation confidence corresponding to the region images; analyze the region images by using the lightweight model to obtain gaze confidence corresponding to the region images; analyze the region images by using the MediaPipe to obtain hand action confidence corresponding to the region images; and perform multi-modal analysis on the person face orientation confidence, the gaze confidence, the hand action confidence, and the region calibration information to obtain the person analysis results; specifically configured to receive preset weight information, wherein the weight information includes face orientation weight, gaze weight, and hand action weight; and perform multi-modal analysis on the weight information, the person face orientation confidence, the gaze confidence, the hand action confidence, and the region calibration information to obtain the person analysis results, wherein the person analysis results include single-camera attention confidence, and the calculation formula is as follows: ; Wherein, C is a single camera on the comprehensive confidence of the character attention to the goods; is a face orientation confidence; is a line of sight confidence; is a hand action confidence; W1, W2, W3 are face orientation weight, line of sight weight and hand action weight, respectively; The person behavior analysis fusion module (205) is configured to fuse the person analysis results to obtain person analysis comprehensive results, specifically configured to fuse the person analysis results to obtain the person analysis comprehensive results, wherein the person analysis comprehensive results include final comprehensive confidence, and the calculation formula is as follows: ; wherein, is the final comprehensive confidence, representing the probability of the character's attention to a certain item or area; Ci is the single-camera confidence of the i-th camera; Wi is the weight of the i-th camera, reflecting the reliability or importance of the camera data; N is the number of cameras participating in fusion; is the sum of the weights of all cameras.
5. A computer device, comprising: A computer program product is provided, which comprises a memory and a processor, wherein the memory stores a computer program capable of being loaded and executed by the processor, and the computer program is capable of executing any one of the methods in claims 1 to 3.
6. A computer-readable storage medium, characterized in that, A computer program product is provided, which comprises a memory and a processor, wherein the memory stores a computer program capable of being loaded and executed by the processor, and the computer program is capable of executing any one of the methods in claims 1 to 3.
Citation Information
Patent Citations
Image processing method and system for intelligent security and protection monitoring
CN118887622A
Information processing device, information processing method, and program
JP2025033544A