Large-scale scene character behavior analysis method, system and device and storage medium
Through the multi-channel long-range camera combined with time synchronization and multimodal analysis, the threshold and area calibration are dynamically adjusted, and the problem of single camera data being affected by viewing angle occlusion in large scenarios is solved, achieving high accuracy and robust character behavior analysis.
Patent Information
- Application Number
- CN202511031897.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-07-25
AI Technical Summary
In the prior art, when analyzing the behavior of characters in large scenarios, single camera data is easily affected by viewing angle occlusion, and cannot accurately capture subtle behavioral clues of small or densely placed items, resulting in poor analysis accuracy and lack of adaptability to dynamic changes in the scene.
Multi-channel vision cameras are used to perform scene area calibration, combining time synchronization and multi-modal analysis, using pedestrian re-identification model to track character trajectories, dynamically update area calibration with semantic segmentation technology, introducing adaptive threshold adjustment mechanism and multi-modal data fusion to optimize multi-source data integration.
It significantly improves the accuracy and robustness of small items' attention analysis, reduces maintenance costs, adapts to changes in different scenarios and item types, and improves the accuracy and efficiency of character behavior analysis in large scenarios.
Smart Images

Figure CN120525971A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and in particular to a method, system, device and storage medium for analyzing human behavior in large-scale scenes. Background Art
[0002] With the increasing informatization and digitization of society, the demand for analyzing human behavior in large-scale scenarios (such as public places, transportation hubs, and industrial production sites) is growing. Existing technologies typically analyze human behavior based on trajectory. For example, in large-scale scenarios (such as furniture stores and auto dealerships), surveillance cameras record people's movements and dwell time to infer their attention. These systems employ simple motion detection or object tracking algorithms, such as those used by Hikvision and Dahua for customer flow analysis.
[0003] However, existing technologies mostly rely on single-camera data and are easily affected by perspective occlusion. Therefore, the attention analysis effect on small or densely placed objects (such as cosmetics and electronic products) is poor and the accuracy is low. It is unable to capture subtle behavioral clues such as gaze direction or body orientation, resulting in poor accuracy of human behavior analysis in large scenes. Summary of the Invention
[0004] In order to improve the accuracy of character behavior analysis in large-scale scenes, the present application provides a large-scale character behavior analysis method, system, device and storage medium.
[0005] In the first aspect, the present application provides a method for analyzing the behavior of people in large-scale scenes, which adopts the following technical solutions: Performing scene area calibration on the plurality of multi-channel long-range cameras and recording a plurality of area calibration information, the area calibration information including coverage areas of the plurality of long-range cameras and information and coordinates of objects in the coverage areas; Using the plurality of multi-channel long-range cameras to collect a plurality of scene data, and time-synchronizing the plurality of scene data; Tracking person trajectories in the scene data using a person re-identification model to obtain a plurality of area stay times and a plurality of area images corresponding to the plurality of area stay times; Performing multimodal analysis on the plurality of regional images to obtain a plurality of character analysis results; A comprehensive character analysis result is obtained by fusing several of the character analysis results.
[0006] Through the above technical solution, this application uses multiple long-range cameras to perform behavioral analysis on people in large scenes, combines time synchronization, optimizes multi-source data integration, significantly improves the accuracy of small object attention analysis, and improves analysis robustness through multi-camera fusion.
[0007] In a specific possible implementation scheme, the method is also based on a semantic segmentation model, and after calibrating the scene area of the plurality of multi-channel long-range cameras and recording a plurality of area calibration information, further includes: Use the semantic segmentation model to identify the boundaries of scene objects; Determining whether the boundary of the item changes; If the boundary of the object changes, the region calibration of the plurality of multi-channel long-range cameras is updated, and the updated region calibration information is recorded.
[0008] Through the above technical solution, this invention achieves automated calibration, reducing maintenance costs. Incorporating semantic segmentation technology, it dynamically adapts to shelf movement or item adjustments, improving system flexibility and accuracy. This covers all dynamic area calibration methods based on computer vision (such as semantic segmentation and object detection) and their applications in behavioral analysis.
[0009] In a specific possible implementation scheme, the step of using a person re-identification model to track person trajectories in the plurality of scene data to obtain a plurality of area stay times and a plurality of area images corresponding to the plurality of area stay times includes: Introducing an adaptive threshold adjustment mechanism to adjust the dwell time threshold according to the scene data; Using a person re-identification model in the scene data to track person trajectories and obtain a plurality of area dwell times; Determining whether the residence time in the plurality of areas exceeds the residence time threshold; If the stay times of the plurality of regions exceed the stay time threshold, the stay times of the plurality of regions and the plurality of region images corresponding to the stay times of the plurality of regions are recorded.
[0010] In a specific implementation scheme, the introduction of an adaptive threshold adjustment mechanism to adjust the dwell time threshold according to the scenario includes: Analyzing some of the scene data to obtain passenger flow data; The dwell time threshold is obtained by adjusting the passenger flow data, and the calculation formula is: ; Among them, T is the residence time threshold, T0 is the basic threshold, D is the crowd density, and α is the adjustment coefficient.
[0011] Through the above technical solution, the present invention improves the robustness and accuracy of behavior analysis by dynamically adjusting thresholds. By incorporating crowd density as a dynamic variable and combining it with mathematical formulas to achieve automated adjustment, this approach outperforms existing static thresholding methods. This encompasses all algorithms and applications that dynamically adjust dwell time thresholds based on crowd density or other scene parameters (such as light and time period).
[0012] In a specific embodiment, the method is further based on a posture estimation model, a lightweight model, and MediaPipe, and the multimodal analysis of the plurality of regional images to obtain the plurality of character analysis results includes: Analyzing the plurality of regional images using the posture estimation model to obtain a plurality of person face orientation confidences corresponding to the plurality of regional images; Analyzing the regional image using the lightweight model to obtain a plurality of sight confidences corresponding to the plurality of regional images; Analyzing the regional image using the MediaPipe to obtain a plurality of hand movement confidences corresponding to the plurality of regional images; A multimodal analysis is performed based on the plurality of character facial orientation confidences, the plurality of line of sight confidences, the plurality of hand movement confidences and the plurality of area calibration information to obtain a plurality of character analysis results.
[0013] In a specific possible implementation scheme, performing multimodal analysis based on the plurality of character facial orientation confidences, the plurality of line of sight confidences, the plurality of hand movement confidences, and the plurality of region calibration information to obtain the plurality of character analysis results includes: Receiving a plurality of preset weight information, wherein the weight information includes a facial orientation weight, a sight weight, and a hand movement weight; A multimodal analysis is performed based on the weight information, the facial orientation confidences, the gaze confidences, the hand movement confidences, and the region calibration information to obtain a plurality of character analysis results, wherein the character analysis results include a single camera focus confidence, and the calculation formula is: ; Among them, C is the comprehensive confidence of a single camera that a person is paying attention to an object; is the face orientation confidence; Confidence of sight; is the hand movement confidence; W1, W2, and W3 are the face orientation weight, gaze weight, and hand movement weight, respectively.
[0014] Through the above technical solution, the present invention integrates multimodal data to improve the accuracy of small item attention analysis (up to 90%). It integrates three behavioral clues and supports dynamic weight adjustment to adapt to different scenarios and item types, which is superior to single-modal analysis.
[0015] In a specific embodiment, fusing several character analysis results to obtain a comprehensive character analysis result includes: A comprehensive character analysis result is obtained by fusing several of the character analysis results. The comprehensive character analysis result includes a final comprehensive confidence score, which is calculated as follows: ; in, is the final comprehensive confidence, which indicates the probability that a person pays attention to a specific object or area; Ci is the single camera confidence of the i-th camera; Wi is the weight of the i-th camera, which reflects the reliability or importance of the camera data; N is the number of cameras involved in the fusion; is the sum of all camera weights.
[0016] Through the above technical solution, this invention improves analysis robustness through multi-camera fusion. By combining time synchronization and viewpoint weighting, it optimizes multi-source data integration and significantly enhances the accuracy of small object attention analysis. This approach encompasses all behavioral analysis methods based on multi-camera data fusion, including time synchronization and confidence weighting mechanisms.
[0017] In a second aspect, the present application provides a large-scale scene character behavior analysis system, which adopts the following technical solution: the system includes: A scene area calibration module, configured to perform scene area calibration on the plurality of the multi-channel long-range cameras and record a plurality of area calibration information, wherein the area calibration information includes coverage areas of the plurality of long-range cameras and information and coordinates of objects in the coverage areas; A scene data acquisition module, configured to acquire a plurality of scene data using the plurality of multi-channel long-range cameras and time-synchronize the plurality of scene data; A person trajectory tracking module is used to track person trajectories in the scene data using a person re-identification model to obtain a plurality of area stay times and a plurality of area images corresponding to the plurality of area stay times; A character behavior analysis module, configured to perform multimodal analysis on the plurality of regional images to obtain a plurality of character analysis results; The character behavior analysis fusion module is used to fuse several character analysis results to obtain a comprehensive character analysis result.
[0018] In a third aspect, the present application provides a computer device that adopts the following technical solution: it includes a memory and a processor, and the memory stores a computer program that can be loaded by the processor and executed by the above-mentioned large-scale scene character behavior analysis method.
[0019] In a fourth aspect, the present application provides a computer-readable storage medium, which adopts the following technical solution: storing a computer program that can be loaded by a processor and execute the above-mentioned large-scale scene character behavior analysis method.
[0020] In summary, this application has the following beneficial technical effects: (1) This application proposes a complete behavior analysis system that integrates pedestrian re-identification, posture estimation, multimodal analysis and dynamic calibration, realizes the integration of multiple technologies, forms a complete behavior analysis chain, uses multiple long-range cameras to analyze the behavior of people in large scenes, combines time synchronization, optimizes multi-source data integration, significantly improves the accuracy of small object attention analysis, and improves the analysis robustness through multi-camera fusion. (2) This invention achieves automated calibration and reduces maintenance costs. Combined with semantic segmentation technology, it dynamically adapts to shelf movement or item adjustments, improving system flexibility and accuracy, and enhancing the accuracy of human behavior analysis in large-scale scenarios.
[0021] (3) This invention improves the robustness and accuracy of behavioral analysis by dynamically adjusting thresholds. By introducing crowd density as a dynamic variable and combining it with mathematical formulas to achieve automated adjustment, this method outperforms existing static threshold methods. This approach encompasses all algorithms and applications that dynamically adjust dwell time thresholds based on crowd density or other scene parameters (e.g., light, time period), improving the accuracy of human behavior analysis in large-scale scenarios.
[0022] (4) This invention integrates multimodal data to improve the accuracy of small-item attention analysis (up to 90%). It integrates three behavioral cues and supports dynamic weight adjustment to adapt to different scenarios and item types. This method is superior to single-modal analysis and improves the accuracy of human behavior analysis in large-scale scenarios.
[0023] (5) Improve the efficiency and data value of business scenarios through automated processes and low-cost implementation (leveraging existing monitoring systems). BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 This is a flow chart of a method for analyzing character behavior in a large scene in an embodiment of the present application.
[0025] Figure 2 This is a structural block diagram of a large-scale scene character behavior analysis method in an embodiment of the present application.
[0026] Reference numerals: 201, scene area calibration module; 202, scene data acquisition module; 203, character trajectory tracking module; 204, character behavior analysis module; 205, character behavior analysis fusion module. DETAILED DESCRIPTION
[0027] The following is combined with Figure 1-Figure 2 This application is described in further detail.
[0028] The embodiment of the present application discloses a method for analyzing human behavior in large-scale scenes, which is used to improve the accuracy of human behavior analysis in large-scale scenes.
[0029] With the increasing informatization and digitization of society, the demand for analyzing human behavior in large-scale scenarios (such as public places, transportation hubs, and industrial production sites) is growing. Existing technologies typically analyze human behavior based on trajectory. For example, in large-scale scenarios (such as furniture stores and auto dealerships), surveillance cameras record people's movements and dwell time to infer their attention. These systems employ simple motion detection or object tracking algorithms, such as those used by Hikvision and Dahua for customer flow analysis.
[0030] However, existing technologies mostly rely on single-camera data and are easily affected by perspective occlusion. Therefore, they have poor analysis effect on attention of small or densely placed objects (such as cosmetics and electronic products), low accuracy, and cannot capture subtle behavioral clues such as gaze direction or body orientation. In addition, they are also very limited, relying on fixed threshold dwell time judgment, lacking adaptability to dynamic changes in the scene, and being easily affected by crowd density or changes in light, resulting in poor accuracy in human behavior analysis in large scenes.
[0031] Therefore, this application proposes a large-scale scene character behavior analysis method, which uses this method to improve the accuracy of character behavior analysis in large-scale scenes.
[0032] The hardware part of this application includes: Multi-channel long-range cameras: Cover different areas within the scene, pre-calibrate object coordinates, and support dynamic area adjustment.
[0033] Processing server: Equipped with high-performance GPU or other inference devices to run face recognition, pedestrian re-identification, pose estimation and multimodal attention analysis models.
[0034] Backend database: stores item information, attention records and behavior analysis results.
[0035] like Figure 1 As shown, the method includes: S10 , calibrating scene areas for multiple long-range cameras and recording a plurality of area calibration information, where the area calibration information includes coverage areas of the multiple long-range cameras and information and coordinates of objects in the coverage areas.
[0036] Specifically, multiple multi-channel long-range cameras are set up in a large scene. Each multi-channel long-range camera covers a portion of the scene, and each area contains corresponding items. For example, in a retail store, the area covered by each multi-channel long-range camera includes shelves and merchandise. First, the scene area must be calibrated for the multi-channel long-range cameras, and the area calibration information is recorded. The area calibration information includes the coverage area of the multi-channel long-range camera and the item information and item coordinates within the coverage area.
[0037] S20, using multiple long-range cameras to collect multiple scene data, and time-synchronizing the multiple scene data.
[0038] Specifically, each multi-channel long-range camera collects scene data of the corresponding area. Before analyzing the collected scene data, the scene data must be time-synchronized.
[0039] S30, using a person re-identification model to track person trajectories in a plurality of scene data to obtain a plurality of area stay times and a plurality of area images corresponding to the plurality of area stay times.
[0040] Specifically, the collected scene data are processed, and a residence time threshold is pre-set. According to the residence time threshold, a pedestrian re-identification model (such as DeepSORT or FairMOT) is used to track the trajectory of the person to obtain the residence time of several areas and several area images corresponding to the residence time of several areas.
[0041] S40, performing multimodal analysis on the plurality of regional images to obtain a plurality of character analysis results.
[0042] Specifically, multimodal attention analysis is introduced to combine facial orientation, gaze estimation (based on eye key points) and hand movements (detecting hand pointing or touching movements), and multimodal analysis is performed on several regional images to obtain several character analysis results.
[0043] S50, integrating several character analysis results to obtain a comprehensive character analysis result.
[0044] Specifically, the coverage areas of different multi-channel long-range cameras may overlap, and the same stay time may be captured by different multi-channel long-range cameras to obtain regional images from different perspectives. Therefore, the character analysis results of multiple cameras in the same time period are fused to obtain a comprehensive character analysis result.
[0045] This application uses multiple long-range cameras to analyze human behavior in large-scale scenes, combines time synchronization, optimizes multi-source data integration, significantly improves the accuracy of small object attention analysis, and improves analysis robustness through multi-camera fusion.
[0046] In one embodiment, in order to improve the accuracy of human behavior analysis in large scenes, after calibrating scene areas for multiple long-range cameras and recording the area calibration information, the following steps may be performed: First, a semantic segmentation model is used to identify the boundaries of objects in the scene. Specifically, a semantic segmentation model is used to identify the boundaries of objects in the scene. For example, a semantic segmentation model (such as DeepLabv3) is used in a retail store scene to identify shelf and item boundaries in real time. Then, it is determined whether the boundaries of the scene objects have changed. If so, the regional calibration of several multi-channel long-range cameras is updated, and the updated regional calibration information is recorded. Specifically, it is determined whether the boundaries of the objects have changed. Taking the retail store scene as an example, the boundaries of the scene objects will change when the shelves move or buyers take away the goods on the shelves. In order to improve the accuracy of human behavior analysis in large scenes, dynamic regional calibration is introduced. That is, when changes in the boundaries of scene objects are detected, the regional calibration information of the multi-channel long-range cameras is updated regularly. Taking the retail store scene as an example, when changes in the boundaries of shelves and objects are identified in real time, the regional coordinates are updated every hour, and the error is controlled at ±5cm. The mapping relationship between the region and the object is automatically generated and stored in the database.
[0047] This invention achieves automated calibration, reducing maintenance costs. Incorporating semantic segmentation technology, it dynamically adapts to shelf movement or item adjustments, improving system flexibility and accuracy. It encompasses all dynamic region calibration methods based on computer vision (such as semantic segmentation and object detection) and their applications in behavioral analysis.
[0048] In one embodiment, to improve the accuracy of human behavior analysis in large-scale scenes, the step of using a person re-identification model to track human trajectories in multiple scene data to obtain multiple region dwell times and multiple region images corresponding to the multiple region dwell times can be specifically performed as follows: First, an adaptive threshold adjustment mechanism is introduced to adjust the residence time threshold according to the scene data. Specifically, crowd density is introduced as a dynamic variable, and the residence time threshold is adaptively adjusted according to the scene data.
[0049] Then, a person re-identification model is used to track the trajectory of the person in several scene data to obtain the residence time in several areas. Specifically, a person re-identification model (such as DeepSORT or FairMOT) is used to track the trajectory of the person based on the residence time threshold to obtain the residence time in several areas.
[0050] Next, determine whether the residence time of several areas exceeds the residence time threshold. If the residence time of several areas exceeds the residence time threshold, then record the residence time of several areas and the images of several areas corresponding to the residence time of several areas. Specifically, during the acquisition process, the residence time of each area collected by each multi-channel long-range camera needs to be compared with the residence time threshold. Only the area residence time exceeding the residence time threshold is recorded. After comparing the residence time of all areas collected by all multi-channel long-range cameras, the residence time of several areas and the images of several areas corresponding to the residence time of several areas are obtained.
[0051] In one embodiment, in order to improve the accuracy of human behavior analysis in large-scale scenes, an adaptive threshold adjustment mechanism is introduced to adjust the dwell time threshold according to the scene. This step can be specifically performed as follows: First, we analyze several scene data to obtain passenger flow data. Specifically, in reality, during peak hours, people stay longer, so we first need to analyze the collected scene data to obtain passenger flow data. Then, the dwell time threshold is adjusted according to the passenger flow data, and the calculation formula is: ; Among them, T is the residence time threshold, T0 is the basic threshold, D is the crowd density, and α is the adjustment coefficient; Specifically, the dwell time threshold is dynamically adjusted based on real-time passenger flow statistics. The calculation formula for the dwell time threshold is: ; Among them, T0 is the basic threshold (such as 5 seconds), D is the crowd density (normalized value 0-1), and α is the adjustment coefficient (0.5).
[0052] T: The actual residence time threshold (in seconds), used to determine whether the person stays in a certain area long enough to trigger subsequent behavior analysis (such as posture estimation).
[0053] T0: Basic residence time threshold, set to 5 seconds, represents the default threshold in a standard scenario (without the influence of special crowd density).
[0054] D: Crowd density, the normalized value is between 0 and 1, indicating the density of the crowd in the scene.
[0055] D=0: The scene is almost empty (e.g., a retail store late at night).
[0056] D=1: The scene is very crowded (such as during the peak period of promotion activities).
[0057] Crowd density can be calculated by detecting the number of people in real time through cameras and normalizing the number of people (for example, number of people / maximum capacity).
[0058] α: Adjustment coefficient, set to 0.5, used to control the impact of crowd density on the threshold. The smaller the value, the smaller the threshold adjustment; the larger the value, the larger the adjustment.
[0059] This invention improves the robustness and accuracy of behavioral analysis by dynamically adjusting thresholds. By incorporating crowd density as a dynamic variable and combining it with mathematical formulas to achieve automated adjustments, it outperforms existing static thresholding methods. This approach encompasses all algorithms and applications that dynamically adjust dwell time thresholds based on crowd density or other scene parameters (such as light and time period).
[0060] In one embodiment, in order to improve the accuracy of character behavior analysis in large scenes, the step of performing multimodal analysis on a plurality of regional images to obtain a plurality of character analysis results can be specifically performed as follows: First, a pose estimation model is used to analyze several regional images to obtain the facial orientation confidence scores of several people corresponding to the regional images. Specifically, for regional images that exceed the threshold, the pose estimation model (such as MoveNet) is input to detect the key points of the human body (left and right shoulders, ears, and eyes, a total of 6 points). The facial orientation angle is calculated (with an error of ±5°) and converted into facial orientation confidence scores, which reflect whether the person's face is facing the target object.
[0061] Secondly, a lightweight model is used to analyze the regional images to obtain several gaze confidences corresponding to the regional images. Specifically, the eye key points are analyzed through a lightweight model (such as GazeNet), and the gaze estimation is obtained based on the eye key point analysis. The gaze estimation is converted into gaze confidence, which reflects whether the gaze is pointing to the target object.
[0062] Next, MediaPipe is used to analyze the regional images to obtain several hand movement confidence levels corresponding to the regional images. Specifically, MediaPipe detects regional images to analyze hand movements, analyzes whether the hand points to or touches the target object, and converts the hand movements into hand movement confidence levels, which reflect whether the hand points to or touches the target object.
[0063] Then, a multimodal analysis is performed based on several facial orientation confidences, several gaze confidences, several hand movement confidences, and several regional calibration information to obtain several person analysis results. Specifically, after obtaining the person's facial orientation angle, gaze estimation, and hand movement, it is necessary to combine the regional calibration information (including information about objects in the area) from multiple long-range cameras. This combined analysis can produce a person analysis result, that is, the comprehensive confidence of a single camera that the person is paying attention to a specific object. After analyzing all multiple long-range cameras, several person analysis results are obtained.
[0064] In one embodiment, to improve the accuracy of character behavior analysis in large scenes, the step of performing multimodal analysis based on multiple character facial orientation confidences, multiple eye gaze confidences, multiple hand movement confidences, and multiple region calibration information to obtain multiple character analysis results can be specifically performed as follows: First, a plurality of preset weight information is received, including face orientation weight, gaze weight, and hand movement weight. Specifically, the ratio of the face orientation weight, gaze weight, and hand movement weight for each multi-channel long-range camera is preset based on scene requirements (such as different object types or camera viewing angles); Then, a multimodal analysis is performed based on several weight information, several character face orientation confidences, several gaze confidences, several hand movement confidences, and several area calibration information to obtain several character analysis results. The character analysis results include the single camera attention confidence, and the calculation formula is: ; Among them, C is the comprehensive confidence of a single camera that a person is paying attention to an object; is the face orientation confidence; Confidence of sight; is the hand action confidence; W1, W2, W3 are the face orientation weight, sight weight and hand action weight respectively; Specifically, after obtaining the person's facial orientation angle, gaze estimation, and hand movements, it is necessary to combine the regional calibration information of multiple long-range cameras (including information about objects in the area). This combined analysis can produce a person analysis result, that is, the comprehensive confidence of a single camera in the person's attention to a certain object. After analyzing all multiple long-range cameras to obtain several person analysis results, the calculation formula for the single camera's attention confidence is: ; Where C is the comprehensive confidence of a single camera that a person is paying attention to an object, ranging from 0 to 1. The higher the value, the greater the possibility that the person is paying attention to the object.
[0065] : Facial orientation confidence, based on the facial orientation angle calculated by the human pose estimation model (such as MoveNet), reflects whether the person's face is facing the target object, ranging from 0 to 1.
[0066] : Gaze confidence, which uses a gaze estimation model (such as GazeNet) to analyze eye key points and determine whether the gaze is pointing to the target object, ranging from 0 to 1.
[0067] : Hand movement confidence, which uses a hand movement detection model (such as MediaPipe) to analyze whether the hand is pointing to or touching the target object, ranging from 0 to 1.
[0068] w1, w2, w3: are the weights of facial orientation, gaze, and hand movements, with default values of 0.4, 0.3, and 0.3, respectively, satisfying w1+w2+w3=1, ensuring that the overall confidence C is between 0 and 1.
[0069] Function of the formula: This formula combines the confidence of three modalities (facial orientation, gaze, and hand gestures) through weighted summation to comprehensively determine whether a person is paying attention to a specific object.
[0070] Multimodal fusion: A single modality (such as facial orientation alone) may lead to misjudgment due to angle, lighting, or occlusion. Combining gaze and hand movements can improve the robustness and accuracy of analysis.
[0071] This invention integrates multimodal data to improve the accuracy of small item attention analysis (up to 90%). It integrates three behavioral cues and supports dynamic weight adjustment to adapt to different scenarios and item types, surpassing single-modal analysis.
[0072] In one embodiment, in order to improve the accuracy of character behavior analysis in large-scale scenarios, the step of fusing several character analysis results to obtain a comprehensive character analysis result can be specifically performed as follows: The comprehensive character analysis result is obtained by integrating several character analysis results. The comprehensive character analysis result includes the final comprehensive confidence, and the calculation formula is: ; in, is the final comprehensive confidence, which indicates the probability that a person pays attention to a specific object or area; Ci is the single camera confidence of the i-th camera; Wi is the weight of the i-th camera, which reflects the reliability or importance of the camera data; N is the number of cameras involved in the fusion; is the sum of all camera weights; Specifically, the coverage areas of different multi-channel long-range cameras may overlap, and the same dwell time may be captured by different multi-channel long-range cameras to obtain regional images from different perspectives. Therefore, the character analysis results of multiple cameras in the same period are fused to obtain the comprehensive character analysis results. The comprehensive character analysis results include the final comprehensive confidence level, and the calculation formula for the final comprehensive confidence level is: ; in, : The final comprehensive confidence, which indicates the probability of a person paying attention to a specific item or area, ranges from 0 to 1. The higher the value, the greater the possibility of attention.
[0073] Ci: Single camera confidence of the i-th camera, calculated by multimodal attention analysis (i.e., Ci ), reflecting the possibility that a person is paying attention to an object under the camera's perspective, ranging from 0 to 1.
[0074] Wi: The weight of the i-th camera, reflecting the reliability or importance of the camera data. The weight is determined based on the quality of the camera's view (such as distance, angle, and occlusion), and usually ranges from 0 to 1.
[0075] N: The number of cameras involved in the fusion (for example, 4 distant cameras in a retail scenario).
[0076] : The sum of all camera weights, used for normalization to ensure Between 0 and 1.
[0077] Formula function: This formula fuses the confidence levels Ci of multiple cameras through weighted averaging to generate a more accurate and robust comprehensive confidence level Cfinal.
[0078] Perspective integration: A single camera may have inaccurate confidence due to perspective occlusion, light changes, or distance. The fusion of multi-camera data can make up for the shortcomings.
[0079] Weight adjustment: Different cameras are assigned different importance through Wi-Fi. For example, cameras with close range and no obstructions have higher weights.
[0080] Time synchronization: The formula implicitly requires time synchronization to ensure that all Ci correspond to the behavior of interest in the same time period.
[0081] This method improves analysis robustness through multi-camera fusion. Combining time synchronization and viewpoint weighting, it optimizes multi-source data integration and significantly enhances the accuracy of small object attention analysis. It encompasses all behavioral analysis methods based on multi-camera data fusion, including time synchronization and confidence weighting mechanisms.
[0082] Based on the above method, the embodiment of the present application also discloses a large-scale scene character behavior analysis system. Figure 2 , the system includes the following modules: A scene area calibration module 201 is used to calibrate the scene area of multiple long-range cameras and record a plurality of area calibration information, wherein the area calibration information includes the coverage area of the multiple long-range cameras and the information and coordinates of objects in the coverage area; The scene data acquisition module 202 is used to collect a plurality of scene data using a plurality of multi-channel perspective cameras and time synchronize the plurality of scene data; The person trajectory tracking module 203 is configured to track person trajectories in a plurality of scene data using a person re-identification model to obtain a plurality of area stay times and a plurality of area images corresponding to the plurality of area stay times; A character behavior analysis module 204 is configured to perform multimodal analysis on a plurality of regional images to obtain a plurality of character analysis results; The character behavior analysis fusion module 205 is used to fuse several character analysis results to obtain a comprehensive character analysis result.
[0083] In one embodiment, the person trajectory tracking module 203 is specifically used to introduce an adaptive threshold adjustment mechanism to adjust the residence time threshold according to the scene data; use the pedestrian re-identification model to track the person trajectory in a number of scene data to obtain the residence time of a number of areas; determine whether the residence time of a number of areas exceeds the residence time threshold; if the residence time of a number of areas exceeds the residence time threshold, record the residence time of a number of areas and a number of area images corresponding to the residence time of a number of areas.
[0084] In one embodiment, the character trajectory tracking module 203 is specifically configured to parse a number of scene data to obtain passenger flow data; The dwell time threshold is obtained by adjusting the passenger flow data. The calculation formula is: ; Among them, T is the residence time threshold, T0 is the basic threshold, D is the crowd density, and α is the adjustment coefficient.
[0085] In one embodiment, the character behavior analysis module 204 is specifically used to use a posture estimation model to analyze several regional images to obtain several character facial orientation confidences corresponding to the several regional images; use a lightweight model to analyze regional images to obtain several line of sight confidences corresponding to the several regional images; use MediaPipe to analyze regional images to obtain several hand movement confidences corresponding to the several regional images; perform multimodal analysis based on several character facial orientation confidences, several line of sight confidences, several hand movement confidences and several regional calibration information to obtain several character analysis results.
[0086] In one embodiment, the character behavior analysis module 204 is specifically configured to receive a plurality of preset weight information, the weight information including a facial orientation weight, a gaze weight, and a hand movement weight; Based on several weight information, several character face orientation confidences, several gaze confidences, several hand movement confidences and several area calibration information, a multimodal analysis is performed to obtain several character analysis results. The character analysis results include the single camera attention confidence, and the calculation formula is: ; Among them, C is the comprehensive confidence of a single camera that a person is paying attention to an object; is the face orientation confidence; Confidence of sight; is the hand movement confidence; W1, W2, and W3 are the face orientation weight, gaze weight, and hand movement weight, respectively.
[0087] In one embodiment, the character behavior analysis fusion module 205 is specifically used to fuse several character analysis results to obtain a comprehensive character analysis result. The comprehensive character analysis result includes a final comprehensive confidence score, which is calculated as follows: ; in, is the final comprehensive confidence, which indicates the probability that a person pays attention to a specific object or area; Ci is the single camera confidence of the i-th camera; Wi is the weight of the i-th camera, which reflects the reliability or importance of the camera data; N is the number of cameras involved in the fusion; is the sum of all camera weights.
[0088] The embodiment of the present application also discloses a computer device.
[0089] Specifically, the computer device includes a memory and a processor, and the memory stores a computer program that can be loaded by the processor and execute the above-mentioned large-scale scene character behavior analysis method.
[0090] The embodiment of the present application also discloses a computer-readable storage medium.
[0091] Specifically, the computer-readable storage medium stores a computer program that can be loaded by a processor and executed such as the above-mentioned large-scale scene character behavior analysis method. The computer-readable storage medium includes, for example: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.
[0092] This specific embodiment is merely an explanation of the present invention and is not intended to limit the present invention. After reading this specification, those skilled in the art may make non-creative modifications to this embodiment as needed. However, as long as such modifications are within the scope of the claims of the present invention, they are protected by patent law.
Claims
1. A method for analyzing human behavior in large scenes, characterized in that: The method is based on multiple long-range cameras and includes: Performing scene area calibration on the plurality of multi-channel long-range cameras and recording a plurality of area calibration information, the area calibration information including coverage areas of the plurality of long-range cameras and information and coordinates of objects in the coverage areas; Using the plurality of multi-channel long-range cameras to collect a plurality of scene data, and time-synchronizing the plurality of scene data; Tracking person trajectories in the scene data using a person re-identification model to obtain a plurality of area stay times and a plurality of area images corresponding to the plurality of area stay times; Performing multimodal analysis on the plurality of regional images to obtain a plurality of character analysis results; fusing several of the character analysis results to obtain a comprehensive character analysis result; The step of using a person re-identification model to track person trajectories in the plurality of scene data to obtain a plurality of area stay times and a plurality of area images corresponding to the plurality of area stay times comprises: Introducing an adaptive threshold adjustment mechanism to adjust the dwell time threshold according to the scene data; Using a person re-identification model in the scene data, tracking person trajectories to obtain a plurality of area dwell times; Determining whether the residence time in the plurality of areas exceeds the residence time threshold; If the stay times of the plurality of regions exceed the stay time threshold, the stay times of the plurality of regions and the plurality of region images corresponding to the stay times of the plurality of regions are recorded.
2. The method according to claim 1, characterized in that The method is also based on a semantic segmentation model, and after calibrating the scene area of the plurality of multi-channel long-range cameras and recording a plurality of area calibration information, further includes: Use the semantic segmentation model to identify the boundaries of scene objects; Determining whether the boundary of the scene object changes; If the boundary of the scene object changes, the region calibration of the plurality of multi-channel long-range cameras is updated, and the updated region calibration information is recorded.
3. The method according to claim 2, characterized in that The introduction of the adaptive threshold adjustment mechanism to adjust the dwell time threshold according to the scenario includes: Analyzing some of the scene data to obtain passenger flow data; The dwell time threshold is obtained by adjusting the passenger flow data, and the calculation formula is: ; Among them, T is the residence time threshold, T0 is the basic threshold, D is the crowd density, and α is the adjustment coefficient.
4. The method according to claim 1, characterized in that The method is also based on a posture estimation model, a lightweight model, and MediaPipe, and the multimodal analysis of the plurality of regional images to obtain the plurality of character analysis results includes: Analyzing the plurality of regional images using the posture estimation model to obtain a plurality of person face orientation confidences corresponding to the plurality of regional images; Analyzing the regional image using the lightweight model to obtain a plurality of sight confidences corresponding to the plurality of regional images; Analyzing the regional image using the MediaPipe to obtain a plurality of hand movement confidences corresponding to the plurality of regional images; A multimodal analysis is performed based on the plurality of character facial orientation confidences, the plurality of line of sight confidences, the plurality of hand movement confidences and the plurality of area calibration information to obtain a plurality of character analysis results.
5. The method according to claim 4, characterized in that: The step of performing multimodal analysis based on the plurality of facial orientation confidences, the plurality of sight confidences, the plurality of hand movement confidences, and the plurality of region calibration information to obtain the plurality of character analysis results includes: Receiving a plurality of preset weight information, wherein the weight information includes a facial orientation weight, a sight weight, and a hand movement weight; A multimodal analysis is performed based on the weight information, the facial orientation confidences, the gaze confidences, the hand movement confidences, and the region calibration information to obtain a plurality of character analysis results, wherein the character analysis results include a single camera focus confidence, and the calculation formula is: ; Among them, C is the comprehensive confidence of a single camera that a person is paying attention to an object; is the face orientation confidence; Confidence of sight; is the hand movement confidence; W1, W2, and W3 are the face orientation weight, gaze weight, and hand movement weight, respectively.
6. The method according to claim 5, characterized in that The fusing of the plurality of character analysis results to obtain a comprehensive character analysis result includes: A comprehensive character analysis result is obtained by fusing several of the character analysis results. The comprehensive character analysis result includes a final comprehensive confidence level, and the calculation formula is: ; in, is the final comprehensive confidence, which indicates the probability that a person pays attention to a specific object or area; Ci is the single camera confidence of the i-th camera; Wi is the weight of the i-th camera, which reflects the reliability or importance of the camera data; N is the number of cameras involved in the fusion; is the sum of all camera weights.
7. A large-scale scene character behavior analysis system, characterized in that: The system is based on several multi-channel long-range cameras and includes: A scene area calibration module (201) is used to calibrate the scene area of the plurality of multi-channel long-range cameras and record a plurality of area calibration information, wherein the area calibration information includes the coverage area of the plurality of multi-channel long-range cameras and the object information and object coordinates in the coverage area; A scene data acquisition module (202) is used to acquire a plurality of scene data using a plurality of the multi-channel long-range cameras and to time-synchronize the plurality of the scene data; The character trajectory tracking module (203) is used to use a pedestrian re-identification model to track the character trajectory in the plurality of scene data to obtain the residence time of the plurality of regions and the plurality of region images corresponding to the residence time of the plurality of regions, and is specifically used to introduce an adaptive threshold adjustment mechanism to adjust the residence time threshold according to the scene data; use the pedestrian re-identification model to track the character trajectory in the plurality of scene data to obtain the residence time of the plurality of regions; determine whether the residence time of the plurality of regions exceeds the residence time threshold; if the residence time of the plurality of regions exceeds the residence time threshold, record the residence time of the plurality of regions and the plurality of region images corresponding to the residence time of the plurality of regions; A character behavior analysis module (204) is used to perform multimodal analysis on the plurality of regional images to obtain a plurality of character analysis results; The character behavior analysis fusion module (205) is used to fuse a number of character analysis results to obtain a comprehensive character analysis result.
8. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and executes the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that A computer program is stored which can be loaded by a processor and execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-camera-based target detection analysis method and device, storage medium, system and robot
CN118485962A
Image processing method and system for intelligent security and protection monitoring
CN118887622A
Deep learning-based behavior analysis camera monitoring method and system
CN120198855A
Method, device and system for real time and multi-camera monitoring of a target object
EP4235570A1
Information processing device, information processing method, and program
JP2025033544A