Multi-camera video crowd analysis system and method based on large model
Through a multi-camera video crowd analysis system based on large models, the problems of data dependence, poor generalization and lack of collaborative processing technology of crowd analysis in the prior art are solved, and the video crowd analysis effect with high accuracy and widely used is achieved.
Patent Information
- Application Number
- CN202510536227.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-12-09
- Filing Date
- 2025-04-27
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the analysis of video population, the existing technology has a large number of problems such as labeling data needs, poor generalization of models, poor universality of different scenarios, gaps in crowd identification and segmentation technology, and lack of unified standards and interfaces for multi-camera video collaborative processing technology.
A multi-camera video crowd analysis system based on large models is adopted, including video preprocessing module, large-model crowd detection and segmentation module, large-model crowd composition analysis module, large-model crowd behavior analysis module, context fusion module and output result visualization module, and population detection, segmentation, composition analysis and behavior analysis are carried out through large models, combined with natural language prompt word design to achieve video understanding and analysis.
It improves the accuracy and application scope of video population analysis, reduces dependence on labeled data, enhances the generalization and universality of the model, and realizes the collaborative processing of multi-camera videos and crowd cross-screen tracking.
Smart Images

Figure CN120088711A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video analysis, and specifically relates to a multi-camera video crowd analysis system and method based on a large model. Background Art
[0002] Deep learning technology is a technology that enables a computer to automatically learn and recognize complex patterns through a large amount of data training, and is widely used in fields such as image recognition, natural language processing, and predictive analysis. The problem with deep learning technology is that it requires a large amount of labeled data, and the generalization of the trained model is not strong. When the test scenario is too different from the training scenario, the test effect will drop sharply.
[0003] The existing invention patent with the publication number "CN105447458A" discloses a large-scale crowd video analysis system and method. The system includes a crowd density calculation module, a crowd foreground segmentation module, a crowd tracking module, a crowd state analysis module, and an event determination module. Among them: the crowd density calculation module, the crowd foreground segmentation module, and the crowd tracking module process video image data to obtain the number of people, the crowd area, and the crowd movement direction and speed respectively; the crowd state analysis module processes and analyzes based on the obtained number of people, crowd area, crowd movement direction and speed, and sends the analysis result to the event determination module; the event determination module is used to determine whether the crowd event is abnormal. This invention proposes multiple different modules, and through the sequential processing of different modules, it realizes the video analysis of a specific scenario and finally outputs the result. Each module uses different visual processing algorithms. For example, the foreground segmentation module uses a fully convolutional network, and the crowd tracking module uses the KLT algorithm, etc. However, this invention also has the following disadvantages:
[0004] 1. Require a large amount of labeled data for training
[0005] a. Strong data dependence: The visual processing algorithms of deep learning usually require a large amount of labeled data for training. These data not only require a large quantity, but also require representativeness to cover various possible situations and changes. However, obtaining and labeling such a data set is often a time-consuming, laborious, and costly task.
[0006] b. Complexity of model training: As the amount of data increases, the computing resources required for training the model also increase exponentially. This includes high-performance computing devices, a large amount of storage space, and long training times. For many research institutions and enterprises, such resource requirements may be unaffordable.
[0007] c. Accuracy of data annotation: The accuracy of data annotation has a crucial impact on the performance of the model. If there are errors or biases in the annotated data, the trained model is likely to fail to correctly learn and understand the inherent laws and characteristics of the data.
[0008] 2. Poor generalization in the same scenario
[0009] Poor generalization is mainly manifested in a significant decline in the performance of the model when the test environment changes. This is usually because the model is too dependent on specific features of the training data (overfitting), and fails to learn the general laws and essential characteristics of the data, or the model is too complex, so that it memorizes the details and noises of the training data. When the distribution of the test data deviates from the training data, or the lighting, background, etc. in the scenario change, the model often fails to make correct predictions and identifications, resulting in a decline in generalization ability.
[0010] 3. Poor versatility in different scenarios
[0011] Poor versatility is mainly manifested in that when the test scenario or task changes, a certain algorithm in the field of vision cannot handle different tasks in this field. Specifically, the existing visual recognition algorithms of deep learning are limited by the training data, and a certain algorithm model can only be applied to a specific scenario. For different task requirements, it is necessary to re-label the data for training, and it is impossible to make a model solve new problems in the field of vision through simple fine-tuning.
[0012] 4. Lack of crowd recognition and segmentation technology
[0013] Existing visual recognition technologies cannot complete the detection and segmentation of crowds in videos. Currently, many algorithms are designed for the detection and segmentation of single individuals, and these algorithms often perform poorly when dealing with multi-person scenarios. They may not be able to accurately distinguish different individuals, or mis-detect multiple individuals as a whole. Lack of global information: Some algorithms may only focus on local information, such as a certain part or feature of a person, while ignoring global information. This results in their poor performance when dealing with complex situations such as occlusion and overlap.
[0014] 5. Lack of collaborative processing technology for multi-camera videos
[0015] The discrete information generated by multiple cameras lacks an effective and fast integration mechanism, making it difficult for users to quickly obtain complete target information. Secondly, due to factors such as perspective changes and lighting differences, the accuracy of collaborative tracking and positioning is challenged, especially in complex scenarios, it is difficult to accurately identify and track targets. The collaborative processing technology of multi-camera videos also faces challenges in standardization and versatility. Currently, there is a lack of unified standards and general-purpose interfaces, which limits the effective collaborative processing between different manufacturers and different models of cameras, and further affects the promotion and application scope of the technology. Summary of the Invention
[0016] The purpose is to provide a multi-camera video crowd analysis system and method based on a large model for the deficiencies of the existing technology, which is used to analyze the composition of the crowd in the image in detail to improve the accuracy and application scope of the analysis. It can identify individuals in the video and track their movements, thereby inferring whether they are acting alone or in groups.
[0017] On the one hand, the present invention provides a multi-camera video crowd analysis system based on a large model, including:
[0018] A video preprocessing module for uniformly encoding the video and sampling the video frames to obtain independent frame images;
[0019] A large model crowd detection and segmentation module for defining crowd detection and segmentation parameters and a crowd prompt word template, and segmenting and labeling the crowd in the frame image based on the large model;
[0020] A large model crowd composition analysis module for defining crowd component type parameters and a component prompt word template, and analyzing the components of the crowd based on the large model according to the results of crowd detection and segmentation to determine the crowd type;
[0021] A large model crowd behavior analysis module for defining behavior type parameters and a behavior prompt word template, tracking the crowd in the main camera video and adjacent camera videos based on the large model according to the results of crowd detection, segmentation and component analysis, outputting a crowd description, and determining the camera ID where the target crowd is most clearly visible, as well as the timestamp and specific location description of the target crowd appearing under this camera;
[0022] A context fusion module for receiving the results of crowd behavior analysis, combining the context information of the video segment, adjusting the camera ID to determine the main camera, and recording the position information of the target crowd under the main camera in chronological order;
[0023] An output result visualization module for presenting the crowd composition and behavior analysis results to the user in a visual form.
[0024] Preferably, the video preprocessing module specifically includes:
[0025] Video reception and transcoding for uniformly encoding and compressing the transmitted video using the FFmpeg toolbox and transcoding it into an MP4 video file using H.264;
[0026] A video frame sampling module for sampling the video according to the set sampling rate to obtain independent frame images;
[0027] Video frame scaling and naming module, which is used to scale the sampled video frames to a specific size proportionally using the OpenCV toolkit, name the video frames according to the time sequence, and record the tags of the video;
[0028] Real-time adaptive enhancement and denoising module, which is used to perform real-time and adaptive enhancement and denoising processing when the video image quality deteriorates.
[0029] Preferably, the large model crowd detection and segmentation module specifically includes:
[0030] Crowd definition module, which is used to define the crowd detection and segmentation parameter as the maximum allowable interval between individuals. If the distance between individuals is less than this distance, they are considered to be a crowd;
[0031] Crowd prompt word template definition module, which is used to define the crowd prompt word template;
[0032] Crowd prompt word output module, which is used to adaptively output the corresponding prompt words according to the different maximum allowable intervals between individuals for the crowd prompt word template;
[0033] Annotation module, which is used to annotate and represent the crowd segmentation results in the form of masks or bounding boxes, and use data annotation tools to add significant annotation boxes into the video content in the video.
[0034] Preferably, the large model crowd composition analysis module specifically includes:
[0035] Component definition module, which is used to define different components and define the component set;
[0036] Component prompt word template definition module, which is used to define the component prompt word template according to the requirement characteristics of component analysis;
[0037] Crowd component type output module, which is used to perform component analysis on the crowd in the frame image by combining the crowd component type parameters and confidence levels, and considering the number of individuals, distance and contact, posture and expression, and environmental factors, and output the component type of each annotation box.
[0038] Preferably, the large model crowd behavior analysis module specifically includes:
[0039] Behavior definition module, which is used to define the behavior types;
[0040] Behavior prompt word template definition module, which is used to define the behavior prompt word template according to the requirement characteristics of behavior analysis;
[0041] Crowd behavior output module, which is used to generate prompt words for crowd recognition, clarity evaluation, and timestamp and location recording based on the input main camera video, adjacent camera video, and crowd segmentation and component analysis results.
[0042] Preferably, the context fusion module specifically includes:
[0043] A camera adjustment module, which is used to automatically select the camera that can most clearly capture the target population according to the position and movement trajectory of the target population identified by the behavior analysis module in the video, and adjust it to the main camera;
[0044] A position information storage module, which is used to record the position information of the target population under the main camera in chronological order, including coordinates, regions, etc.
[0045] Preferably, the output result visualization module specifically includes:
[0046] A data visualization module, which is used to visually display the results of population composition and behavior analysis in the form of charts, images, etc.;
[0047] A report generation module, which is used to generate a detailed analysis report according to user needs.
[0048] On the other hand, the present invention provides a multi-camera video crowd analysis method based on a large model, including the following steps:
[0049] S1. Preprocess the input video segment, uniformly encode the video and perform video frame sampling to obtain independent frame images;
[0050] S2. Set the initial main camera and other adjacent cameras according to the identified target population, set the main camera according to the identified target population and monitoring requirements, and determine other cameras adjacent to the main camera as auxiliary cameras;
[0051] S3. Define crowd detection segmentation parameters and crowd prompt word templates according to the identified target population and actual application scenarios;
[0052] S4. Automatically generate crowd segmentation prompt words based on the current scene and identified target;
[0053] S5. Input the crowd segmentation prompt words and the preprocessed main camera video into the pre-trained large model for crowd segmentation processing, and output the crowd segmentation result;
[0054] S6. Define crowd components and component prompt word templates according to the crowd segmentation result and the identified target population;
[0055] S7. Automatically generate component analysis prompt words based on the definition of crowd components;
[0056] S8. Input the crowd segmentation result and the component analysis prompt words into the large model for crowd component analysis, and output the component analysis result;
[0057] S9. Define a behavior analysis prompt word template based on the component analysis results and the target population, and automatically generate behavior analysis prompt words;
[0058] S10. Input the main camera video, adjacent camera videos, and behavior analysis prompt words into a large model to analyze the crowd behavior and output the results of the behavior analysis;
[0059] S11. Correct and optimize the results of the behavior analysis according to the context information;
[0060] S12. According to the corrected results of the behavior analysis and whether the recognition purpose is achieved or the video ends, determine whether to continue to loop and execute the processes of steps S4 to S11 until the stop condition is met.
[0061] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0062] 1. The present invention uses a trained large model as the benchmark model, and guides the large model to understand different video tasks by designing prompt words and thinking paths, eliminating the process of retraining for different tasks. The large model has stronger generalization ability and is hardly affected by natural factors such as light in the video.
[0063] 2. The present invention uses a large model for video understanding output. The large model can fully understand the content of the entire video, and perform operations such as person extraction, composition analysis, and behavior analysis on the content presented in the video, improving the generalization of visual recognition and analysis of people.
[0064] 3. Through the design of natural language prompt words, the large model is guided to think about a certain scene task, fully invoking the general decision-making ability of the large model, making a decision output for the new scene task. The large model has better versatility. For different tasks, it can be guided by simple natural language instructions. Through the definition of crowd, component, and behavior, the automatic generation of prompt words for three different video recognition tasks is realized.
[0065] 4. A large model video understanding and analysis system is constructed. In the key parts of the system (video segmentation, crowd composition analysis, crowd behavior analysis), the large model technology with stronger versatility and generalization ability is used to solve the problem, improving the robustness and versatility of the entire video understanding and analysis system. The understanding ability is strong. During the full training process of the large model, a variety of data are included. Therefore, in the inference and recognition process, it can consider more comprehensively in an anthropomorphic manner, rather than the traditional feature-based recognition algorithm.
[0066] 5. It can realize cross-screen tracking of crowds with multiple cameras, so as to solve the pain points of traditional manual single-person / crowd tracking in multiple cameras. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 This is the overall process schematic diagram of the present invention;
[0068] Figure 2 This is the structural diagram of a multi-camera video crowd analysis device based on a large model provided by the present invention;
[0069] Figure 3 This is the schematic diagram of large model video segmentation provided by the present invention;
[0070] Figure 4 This is the flow chart of large model crowd component analysis provided by the present invention;
[0071] Figure 5 This is the flow chart of crowd behavior analysis provided by the present invention. Specific embodiments
[0072] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts belong to the scope of protection of the present invention.
[0073] As an important branch in the field of deep learning, large model technology refers to huge deep learning models with billions or even tens of billions of parameters. Compared with traditional deep learning models, this technology has the following remarkable characteristics and advantages:
[0074] 1. Ultra-large scale parameters: Large models have a huge number of parameters, far exceeding traditional deep learning models. These parameters enable large models to capture and understand the internal laws and patterns of data more accurately, providing powerful representation capabilities.
[0075] 2. Learning ability and generalization ability: With ultra-large scale parameters and complex network structures, large models show powerful learning abilities, capable of processing massive amounts of data and learning complex features therein. At the same time, large models usually have better generalization abilities and can perform well on new and unseen data.
[0076] Through the training of massive amounts of data and the interactive learning of multi-modal knowledge, large models can automatically extract complex features and patterns in images and videos, achieving a deep understanding of the content. In the field of crowd analysis, traditional analysis methods mostly rely on simple behavior pattern recognition, and their accuracy has been greatly challenged in complex scenarios (such as crowded crowds, severe occlusions, lighting changes, etc.).
[0077] Such as Figure 1As shown in the figure, the present invention provides a multi-camera video crowd analysis system based on a large model. In order to achieve efficient detection and segmentation of specific crowds in multi-camera video clips, the present invention first designs a prompt word generation module for crowd definition and segmentation. This module can automatically generate accurate prompt words ( ). Then, these prompt words are input into a pre-trained large model together with the main camera video to obtain accurate crowd segmentation results. On this basis, the present invention further conducts a detailed component analysis on the segmented crowd through a component analysis module. This module also uses prompt word generation technology to generate prompt words related to component analysis ( ), to assist the large model in outputting detailed component analysis results. In the behavior analysis stage, the present invention combines the component analysis results and the video content of other adjacent cameras, and generates corresponding prompt words through a behavior analysis prompt word module ( ), so as to comprehensively capture and understand the behavior patterns of the target crowd. Finally, the present invention can not only determine the exact position of a certain crowd in a specific camera at a certain time point (through the recognition frame coordinates), but also deeply understand its behavior characteristics and its interaction with the surrounding environment. The modules included in the system of the present invention and the implementation process are specifically as follows:
[0078] 1. Video preprocessing module
[0079] The video clip to be detected enters the video preprocessing module. The main functions of the video preprocessing module are as follows:
[0080] (1) Video reception and transcoding:
[0081] Use the FFmpeg toolbox or other similar tools to uniformly transcode the transmitted video into an MP4 video file by H.264 encoding compression. This step ensures the unity and compatibility of the video format and provides convenience for subsequent processing.
[0082] (2) Video frame sampling:
[0083] Sample the video according to the set sampling rate (such as sampling 1 frame every 2 frames) to obtain independent frame images. The selection of the sampling rate should be based on the characteristics of the video content and processing requirements to balance processing efficiency and information loss.
[0084] (3) Video frame scaling and naming:
[0085] Use the OpenCV toolkit or other image processing libraries to scale the sampled video frames proportionally to a specific size (such as an image with a width of 256). Name the video frames according to the time sequence and record the label of the video. Scaling and naming are helpful for subsequent feature extraction and data processing.
[0086] (4) Real-time Adaptive Enhancement and Denoising:
[0087] When the quality of video images deteriorates (such as in foggy days, nights, etc.), perform real-time and adaptive enhancement and denoising processing. This includes enhancement algorithms based on histogram equalization, denoising techniques based on non-local mean algorithms, etc.
[0088] 2. Large Model Crowd Detection and Segmentation Module
[0089] The large model crowd detection and segmentation module is the core part of the video analysis system, responsible for automatically identifying and segmenting the crowd from video frames.
[0090] The first step: Define the crowd
[0091] Define the crowd detection and segmentation parameters, and the maximum allowable interval between individuals , if the distance between individuals is less than this distance, they are considered to be a crowd.
[0092] The second step: Define the prompt template
[0093] Then define the prompt template for the crowd detection and segmentation module as follows:
[0094] (1) Task description:
[0095] "Please perform crowd segmentation detection on the provided video. Note that what is segmented is the crowd. When a single person is significantly far away from others, label it as a single-person crowd."
[0096] (2) Input data description:
[0097] "The input is a video containing crowd activities. The video format is MP4, the resolution is 1920x1080, and the frame rate is 30fps."
[0098] (3) Specific guidance or parameters:
[0099] " (Maximum allowable interval parameter): When performing crowd segmentation, please consider the distance between individuals. If the distance between two individuals is less than or equal to , then they are considered to belong to the same crowd.
[0100] Real-time (optional): If possible, while maintaining accuracy, try to improve the processing speed to meet the real-time requirements."
[0101] (4) Output requirements:
[0102] "The output should be the result of crowd segmentation, which can be represented in the form of a mask or bounding boxes. Each detected crowd area should be clearly marked, and different crowd areas should be avoided from being wrongly merged together."
[0103] (5)Example or template (optional):
[0104] Example videos and corresponding segmentation results can be provided for reference. This is common knowledge in the art and will not be specifically described in this invention.
[0105] (6)Context information (optional):
[0106] The specific background information or scene included in the video can be described. This is common knowledge in the art and will not be specifically described in this invention
[0107] The third step: Prompt generation
[0108] For the maximum allowed interval between different individuals, the prompt template adaptively outputs the corresponding prompt. Under this template, the input parameter is the maximum allowed interval between individuals, and the output content is the corresponding prompt , and the specific formula can be expressed as:
[0109]
[0110] Module example
[0111] Assume the setting value is 1, and a prompt example is output as follows:
[0112] "Please perform crowd segmentation detection on the provided video in MP4 format, with a resolution of 1920x1080 and a frame rate of 30fps. When processing, please use value 1 to determine which individuals belong to the same crowd. If the distance between two individuals is less than or equal to 1 unit, they should be considered part of the same crowd. Please output the result of crowd segmentation, and you can use bounding boxes or other appropriate representations to mark each crowd area. It can be represented in the form of a mask or bounding boxes. Please try to maintain accuracy and, if possible, improve the processing speed to meet real-time requirements."
[0113] After being processed by the large model, the example result output by the large model is as follows (assuming the resolution of the video frame is 1920x1080, and the detected crowd consists of individuals A, B, and C, who together form a labeled box):
[0114] Crowd annotation box coordinates (top - left and bottom - right): The example is (x_min, y_min), (x_max, y_max). After detection and segmentation, the result is: The annotation box coordinates of crowd 1 are (200, 300), (770, 870), and this crowd is composed of 3 people.
[0115] Step 4: Adding annotation results. For the obtained segmentation results, use a data annotation tool to add prominent annotation boxes into the video content in the video for subsequent component analysis.
[0116] 3. Large - model crowd composition analysis module
[0117] This module is mainly responsible for analyzing the crowd composition in the video. According to the results of crowd detection, it analyzes the specified crowd, obtains the composition of the crowd, determines whether it belongs to a certain defined crowd set, and outputs component types, confidence levels, positions, descriptions, etc.
[0118] Step 1: Definition of component types
[0119] First, define different components and a component set , which mainly includes
[0120] Single person : Only contains one individual.
[0121] Couple : Usually composed of two individuals, with an intimate relationship, often maintaining a close distance and contact.
[0122] Friends : Composed of two or more individuals, with a friendly relationship, but not necessarily as intimate as a couple.
[0123] Family : Composed of two or more individuals, usually including members of different age groups, such as parents and children, or grandparents, parents, and children, etc.
[0124] Step 2: Definition of prompt - word templates
[0125] According to the requirement characteristics of component analysis, define the prompt - word templates as follows:
[0126] (1)Overall task description
[0127] "Please perform component analysis on the crowd in the given video frame or image to determine the component types of each crowd area."
[0128] (2)Input data description
[0129] The input data is a video frame or image containing crowds, where each crowd area is represented by a bounding box.
[0130] The bounding boxes have been automatically generated according to the crowd segmentation algorithm.
[0131] (3)Component type parameter
[0132] Please generate corresponding prompt words according to the following component type parameters:
[0133] a. Single Person: The bounding box contains only the body or features of one individual.
[0134] b. Couple: The bounding box contains two individuals with a close relationship, maintaining a relatively close distance and possible physical contact.
[0135] c. Friends: The bounding box contains two or more individuals with a friendly relationship, but not necessarily as close as a couple.
[0136] d. Family: The bounding box contains two or more individuals, usually including members of different ages, such as parents and children.
[0137] (4)Prompt word template
[0138] For each component type parameter, generate the following prompt words:
[0139] a. Single Person prompt words
[0140] Please analyze whether there is only one individual's body or features in the bounding box. If so, classify it as a single person.
[0141] b. Couple prompt words
[0142] Please check whether the bounding box contains two individuals, whether they maintain a relatively close distance and have possible physical contact. If so, classify it as a couple.
[0143] c. Friends prompt words
[0144] Please analyze whether the bounding box contains two or more individuals, whether there are signs of conversation, interaction, etc. between them, but without significant physical contact. If so, classify it as friends.
[0145] d. Family prompt words
[0146] Please check whether the labeled box contains two or more individuals and whether there are members of different age groups between them, such as parents and children. If it meets the family characteristics, classify it as a family.
[0147] (5) Output requirements
[0148] a. For each labeled box, output its component type (selected from the above parameters) and provide a possible confidence level (optional).
[0149] b. The output results should be clear, accurate, and include all detected crowd areas as much as possible.
[0150] (6) Examples or templates (optional)
[0151] A video frame or image containing examples of different component types and the corresponding output results can be provided as reference.
[0152] Step 3: Prompt word generation
[0153] When using a large model to analyze population components, the present invention can use different component types as parameters and generate different prompt words based on these parameters. The specific formula can be expressed as:
[0154]
[0155] in For the component type Components.
[0156] Module example:
[0157] Here are some examples of prompt words for "couples":
[0158] Please perform a component analysis on the crowd in a given video frame or image, paying special attention to the pairs of individuals that may constitute a couple within the annotated box. When performing the analysis, please pay attention to the following points:
[0159] (1) Number of individuals: Couples usually consist of two individuals, so look for areas where the labeled boxes contain only two individuals.
[0160] (2) Distance and contact: Couples usually maintain a close distance and may have physical contact, such as holding hands, hugging, etc. Please check whether the two individuals in the marked box meet these conditions.
[0161] (3) Posture and expression: Couples usually show intimacy and tacit understanding through their posture and expression. Please observe the posture and expression of the two individuals in the marked box to see if they show the typical characteristics of a couple.
[0162] (4) Environmental factors: Couples often appear in romantic places or environments, such as park benches, restaurant corners, etc. Consider whether the background environment in the annotation box supports the presence of couples.
[0163] Based on the above analysis, if the labeled box contains only two individuals, they are close to each other and have physical contact, their postures and expressions also show the characteristics of intimacy and tacit understanding, and the background environment also supports the existence of a couple, then the labeled box can be classified as a couple component. Please output the component type of each labeled box. If it meets the characteristics of a couple, it is marked as "couple" and a possible confidence level is provided (optional). When conducting the analysis, consider multiple factors to ensure the accuracy of the classification. At the same time, pay attention to those labeled boxes that may contain multiple individuals but do not meet the characteristics of a couple.
[0164] The above prompt words are input into the big model, and the big model output example is as follows: [
[0166] {
[0167] "Marker box ID": "1",
[0168] "Component Type": "Couple",
[0169] "Confidence": 0.95,
[0170] "Position": [(100,100), (400,700)],
[0171] "Feature description": "The marked box contains two individuals, who are close to each other and holding hands. Their posture and expression show intimacy and tacit understanding. The background environment is a park bench, which is a typical scene for couples."
[0172] }, ]
[0174] 4. Large model crowd behavior analysis module
[0175] This module focuses on the behavior patterns and trends of a given group of people in a video. It uses a large model to analyze the behavior of people in a video frame, identify various behavior patterns, and predict and interpret behavior trends based on them.
[0176] Step 1: Behavior Type Definition
[0177] First, the present invention needs to define a series of possible behavior types These behavior types should be determined based on specific application scenarios and requirements. Here are some example behavior types:
[0178] Gather (Gathering): People gather together, perhaps engaged in some activity or waiting for something.
[0179] Walking (Walking): People are moving, perhaps walking, hurrying, or fleeing.
[0180] Conflict (Conflict): Disputes, fights, or other forms of conflict occur among the people.
[0181] Chasing (Chasing): One or more individuals are chasing another or more individuals.
[0182] ... (Other behavior types)
[0183] Step 2: Definition of the prompt template
[0184] Define the prompt template according to the requirement characteristics of behavior analysis. As follows:
[0185] (1) Overall task description
[0186] This task aims to conduct behavior analysis and crowd tracking on the input main camera video and adjacent camera videos. By identifying the specified crowd and tracking its movement trajectory under different cameras, finally determine which camera the crowd is clearest in and output its specific location and time information under that camera.
[0187] (2) Input data description
[0188] Video data:
[0189] Main camera video: The main video file to be analyzed, containing the core or important perspectives.
[0190] Adjacent camera video: The video captured by other cameras adjacent to the main camera, used to provide more comprehensive scene information.
[0191] (3) Type parameters
[0192] Behavior type: An optional parameter used to specify the specific behavior type to be analyzed (such as walking, gathering, etc.). In this template, the behavior type is not specific, but focuses on crowd tracking and clarity evaluation.
[0193] Crowd characteristics: Feature descriptions used to identify the target crowd, such as clothing, quantity, behavior patterns, etc.
[0194] (4) Output type
[0195] Behavior Analysis Report: A report containing the behavior description of the target population, the occurrence time, the ID of the clearest camera, the specific location, and the remarks information.
[0196] (5)Prompt Template
[0197] Input Parameters:
[0198] Specify the ID of the main camera video to be analyzed.
[0199] The list of adjacent camera videos and their corresponding IDs.
[0200] The characteristic description of the target population.
[0201] Output Content:
[0202] Behavior Analysis Report, including the following content:
[0203] Population Description: A brief description of the target population, such as quantity, characteristics, etc.
[0204] Behavior Description (Optional): If the behavior type is analyzed, it includes the description of behavior characteristics.
[0205] Timestamp: The time when the target population appears under the clearest camera.
[0206] ID of the Clearest Camera: The ID of the camera where the target population is most clearly visible.
[0207] Location Information: The specific location of the target population under the clearest camera (using xy coordinates or descriptive location).
[0208] Remarks: Other information or observations that need to be explained.
[0209] Step 3: Prompt Generation
[0210] Input the specified population to be identified and analyzed, and the corresponding prompt analysis module generates different behavior analysis prompts . The specific formula can be expressed as:
[0211]
[0212] Module Example
[0213] Suppose the population to be analyzed is Population 1 (couple), then the example of the output prompt is:
[0214] Task Description:
[0215] Please analyze the input main camera video and the adjacent camera video, identify the people group 1 (a couple) in the videos, track their performances under different cameras, determine which camera shows them most clearly, and finally output the time when this couple is at what position (specific xy coordinates or descriptive location) under which camera (camera ID).
[0216] Thought process:
[0217] (1) People group identification:
[0218] Please identify the target people group - the couple (people group 1) in the input main camera video and the adjacent camera video. A couple usually shows as two people closely accompanying each other and may have intimate actions.
[0219] (2) Behavior tracking:
[0220] Once the couple (people group 1) is identified, please track their movement trajectories in the video, especially pay attention to their transfer between different cameras.
[0221] (3) Clarity evaluation:
[0222] Analyze the clarity of the couple (people group 1) in each camera video. This includes evaluating the resolution, contrast, brightness of the image, and the contrast between the couple and the background.
[0223] Determine under which camera the couple (people group 1) is shown most clearly.
[0224] (4) Timestamp and position recording:
[0225] Record the timestamp when the couple (people group 1) appears in the clearest camera.
[0226] Provide the specific position of the couple (people group 1) under this camera. If the video provides an xy coordinate system, use xy coordinates; otherwise, use a descriptive location (such as "in the center of the square", "at the street corner", etc.).
[0227] Expected output:
[0228] Description of the couple (people group 1): Two people are wearing casual clothes, closely accompanying each other, and having intimate actions.
[0229] Timestamp: The specific time when the couple (people group 1) appears in the video (e.g., 2023-10-23T17:30:00Z).
[0230] Camera ID: The camera ID where the couple (people group 1) is most clearly visible (e.g., Cam02).
[0231] Location information: The specific location of the couple (Group 1) under this camera (e.g., xy coordinates are {x: 400, y: 500} or the descriptive location "next to the park bench").
[0232] Input the above prompt into the large model, and the large model outputs examples as follows:
[0233] {
[0234] "Crowd description": "A couple (Group 1), both dressed in casual clothes, walking hand in hand",
[0235] "Timestamp": "2023-10-23T17:30:00Z",
[0236] "Camera ID": "Cam02",
[0237] "Location information": [(100,100), (300,500)],
[0238] "Remarks": "The couple (Group 1) is clearest under the Cam02 camera and is strolling leisurely in the park"
[0239] }
[0240] 5. Context Fusion Module
[0241] The context fusion module is a key component in the video surveillance system. It receives the output results of the behavior analysis module and combines the context information of the video clip to achieve the following functions:
[0242] Automatic adjustment of camera ID: According to the position and movement trajectory of the target crowd identified by the behavior analysis module in the video, automatically select the camera that can most clearly capture the target crowd and adjust it to the main camera. Specifically as follows:
[0243] The behavior analysis module outputs the location information and movement trajectory of the target crowd. The context fusion module evaluates the capture effect of each camera on the target crowd based on this information, including clarity, perspective, occlusion, etc. Select the camera with the best capture effect and adjust it to the main camera for subsequent recording of location information.
[0244] Storage of location information: Record the location information of the target crowd under the main camera in chronological order, including coordinates, regions, etc. Specifically as follows:
[0245] After the main camera is determined, the context fusion module starts to record the location information of the target population under this camera. The location information includes the coordinates of the target population in the video frame (if available), the area where they are located (such as a square, street, etc.), and the timestamp. The location information is stored in chronological order for subsequent query and analysis.
[0246] 6. Output result visualization module
[0247] This module is responsible for presenting the crowd composition and behavior analysis results obtained from the large model analysis to the user in a visual form, so that the user can intuitively understand and apply these results. Specifically, it includes the following visualization methods:
[0248] Data visualization: Visualize the results of crowd composition and behavior analysis in the form of charts, images, etc., such as bar charts, pie charts, heat maps, etc.
[0249] Interactive interface: Includes a user interface that enables users to conveniently view and analyze the results. The interface should include necessary control options and parameter settings to meet the different needs of users.
[0250] Report generation: Generate a detailed analysis report according to the user's needs, including crowd composition analysis, behavior analysis, etc., as well as corresponding visual charts and data tables.
[0251] On the other hand, the present invention also provides a multi-camera video crowd analysis method based on a large model. The specific implementation steps are as follows:
[0252] The first step: Preprocess the input video clip, including operations such as quality adaptive adjustment, noise reduction, and frame rate standardization to optimize the effect of subsequent processing. Specifically, use the FFmpeg tool for video format conversion, compress and transcode all videos into MP4 format using H.264 encoding, then perform frame sampling on the video according to the set sampling rate (sample one frame every two frames), then scale the sampled video frames proportionally, and name them in the order of events; then use an enhancement algorithm based on histogram equalization to denoise the images.
[0253] The second step: According to the recognition target, set the initial main camera and other adjacent cameras. According to the recognition target and monitoring requirements, set the main camera and determine other cameras adjacent to the main camera as auxiliary cameras. The definition of an adjacent camera is a camera that continuously expands the recognition range of the main camera.
[0254] The third step: Define the maximum allowable interval of the crowd according to the recognition target and the actual application scenario to ensure the accuracy and effectiveness of crowd segmentation. Specifically, set different maximum allowable intervals between individuals according to the characteristics of the target population to be recognized. , for example, when there are many people in the video background, set , when the total number of background people in the video is small, it can be set , specifically The value needs to be set considering the actual scenario.
[0255] Step 4: Use the crowd definition and segmentation prompt word generation module to automatically generate crowd segmentation prompt words based on the current scenario and recognition target . Specifically, the prompt word template contains parameter descriptions of the video, such as format, resolution, frame rate, and also contains a description of the task, indicating that the distance between two individuals in the video needs to be calculated, and the value is used to judge whether they belong to the same crowd. In addition, it also needs to contain restrictions on the output format, clearly indicating that the output is in the form of a mask or a bounding box. Based on the task scenario and recognition template, according to the defined template, automatically generate crowd segmentation prompt words .
[0256] Step 5: Input the prompt words and the preprocessed main camera video into the pre-trained large model, and use the powerful capabilities of the large model to perform crowd segmentation processing and output the crowd segmentation results. Specifically, is a natural language expression. Through the word embedding model, the prompt words are encoded to obtain a fixed-length vector; at the same time, the input image data (this usually involves each pixel in the image or the data after transformation) is vectorized to also obtain a fixed-length vector. The two vectors are input into the pre-trained large model, and through the input layer, intermediate layer (Transformer layer) and output layer of the large model, and finally output the specified content according to the requirements, that is, the bounding box positions of each crowd.
[0257] Step 6: Define the components of the crowd according to the crowd segmentation results and recognition target. The components can be based on different features, such as age, gender, behavior, etc. Specifically, the definitions are as follows: single person : Only contains one individual. Couple : Usually consists of two individuals, with an intimate relationship and often maintaining a close distance and contact. Friends : Consists of two or more individuals, with a friendly relationship, but not necessarily as intimate as a couple. Family : Consists of two or more individuals, usually including members of different age groups, such as parents and children, or grandparents, parents and children, etc.
[0258] Step 7: Based on the definition of the crowd components, use the component analysis prompt word generation module to automatically generate prompt words for the large model to perform component analysis Specifically, the component analysis prompt words It should include the number of individuals in the crowd, the distance between individuals and their contact with each other, the postures and expressions of individuals in the crowd, and a simple description of the links. In addition, it also includes format restrictions on the output content of the large model to constrain the output format of the large model.
[0259] Step 8: Segment the crowd and the prompt words The crowd segmentation results are obtained in the fifth step, and the component analysis prompt words are obtained in the seventh step. The two are input into the large model through the input layer, the middle layer (Transformer layer) and the output layer, and finally the result of the component analysis is output. Output as required. For example: [
[0261] {
[0262] "Marker box ID": "1",
[0263] "Component Type": "Couple",
[0264] "Confidence": 0.95,
[0265] "Position": [(100,100), (400,700)],
[0266] "Feature description": "The marked box contains two individuals, who are close to each other and holding hands. Their posture and expression show intimacy and tacit understanding. The background environment is a park bench, which is a typical scene for couples."
[0267] }, ]
[0269] Step 9: Based on the results of component analysis and the target population components, the behavior analysis prompt word module is used to automatically generate behavior analysis prompt words. Specifically, the definition of behavior includes focusing, walking, conflict, chasing, etc. The behavior analysis prompts include the overall task description, the description of the input data, and the restrictions on the output type, which are used to guide the large model to perform reasoning thinking.
[0270] Step 10: Put the main camera video, the adjacent camera video and the behavior analysis prompt words The main camera video information is obtained in step 8, the adjacent camera information is obtained in step 2, and the behavior analysis prompt word is obtained in step 3. It is generated in the ninth step. After vectorizing these three simultaneously and feeding them into the large model, through the processing of the input layer, intermediate layer, and output layer of the large model, the output result is obtained. For example:
[0271] {
[0272] "Crowd description": "Couple (Crowd 1), the two are wearing casual clothes and walking hand in hand",
[0273] "Timestamp": "2023-10-23T17:30:00Z",
[0274] "Camera ID": "Cam02",
[0275] "Location information": [(100,100), (300,500)],
[0276] "Remarks": "The couple (Crowd 1) is the clearest under the Cam02 camera and is strolling leisurely in the park"
[0277] }
[0278] Step 11: The result of the behavior analysis will be input into the context processing module, and the result will be corrected and optimized according to the context information. Specifically, the context processing module receives the output result of the behavior analysis module, and adjusts the encoding of the main camera and the adjacent cameras according to the clearest camera ID given in the analysis result. In addition, the location information of the target crowd under the main camera is cumulatively recorded to record the movement trajectory of the target crowd.
[0279] Step 12: According to the corrected behavior analysis result, and whether the recognition purpose is achieved or the video ends, it is judged whether to continue to loop and execute the processes of Step 4 to Step 11 until the stop condition is met.
[0280] Corresponding to the foregoing embodiment of a large model-based multi-camera video crowd analysis method, the present invention also provides an embodiment of a large model-based multi-camera video crowd analysis device.
[0281] See Figure 2 , an embodiment of a large model-based multi-camera video crowd analysis device provided by the embodiment of the present invention includes a memory and one or more processors, an executable code is stored in the memory, and when the processor executes the executable code, it is used to implement a large model-based multi-camera video crowd analysis method in the above embodiment.
[0282] An embodiment of a multi-camera video crowd analysis device based on a large model provided by the present invention can be applied to any device with data processing capabilities, and such a device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. From a hardware level, as Figure 2 shown, it is a hardware structure diagram of any device with data processing capabilities where the multi-camera video crowd analysis device based on a large model provided by the present invention is located. In addition to Figure 2 the processor, memory, network interface, and non-volatile memory shown, usually according to the actual functions of the device with data processing capabilities where the embodiment device is located, other hardware may also be included, which will not be elaborated here.
[0283] For the specific implementation process of the functions and roles of each unit in the above device, please refer to the implementation process of the corresponding steps in the above method, which will not be elaborated here.
[0284] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present invention solution. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0285] The embodiment of the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements a multi-camera video crowd analysis method based on a large model in the above embodiment.
[0286] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or a memory. The computer-readable storage medium may also be an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc., equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or is to be output.
[0287] The present invention also provides a computer program product, including a computer program, which when executed by a processor, implements the described method for multi-camera video crowd analysis based on a large model.
[0288] Although the present invention has been described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the present invention. Therefore, it should be understood that many modifications can be made to the exemplary embodiments, and other arrangements can be designed, as long as they do not depart from the spirit and scope of the present invention defined by the appended claims. It should be understood that the different dependent claims and the features described herein can be combined in a manner different from that described in the original claims. It should also be understood that the features described in connection with a single embodiment can be used in other described embodiments.
Claims
1. A multi-camera video crowd analysis system based on a large model, characterized in that: include: The video preprocessing module is used to uniformly encode the video and perform video frame sampling to obtain independent frame images; The large-model crowd detection and segmentation module is used to define crowd detection and segmentation parameters and crowd prompt word templates, and segment and annotate the crowd in the frame image based on the large model; The large-model crowd composition analysis module is used to define crowd component type parameters and component prompt word templates, analyze crowd components based on the large model according to the results of crowd detection and segmentation, and determine the crowd type; The large-model crowd behavior analysis module is used to define behavior type parameters and behavior prompt word templates. According to the crowd detection segmentation and component analysis results, the module tracks the crowd on the main camera video and adjacent camera videos based on the large model, outputs the crowd description, and determines the camera ID where the target crowd is most clearly visible, as well as the timestamp and specific location description of the target crowd appearing under the camera; The context fusion module is used to receive the crowd behavior analysis results, and combined with the context information of the video clip, adjust the camera ID to determine the main camera, and record the location information of the target crowd under the main camera in chronological order; The output result visualization module is used to display the crowd composition and behavior analysis results to the user in a visual form.
2. According to the large model-based multi-camera video crowd analysis system of claim 1, it is characterized in that: The video preprocessing module specifically includes: Video reception and transcoding, used to use the FFmpeg toolbox to uniformly compress and transcode the transmitted videos into MP4 video files using H.264 encoding; The video frame sampling module is used to sample the video according to the set sampling rate to obtain independent frame images; The video frame scaling and naming module is used to scale the sampled video frames to a specific size using the OpenCV toolkit, name the video frames according to the time sequence, and record the labels of the video; The real-time adaptive enhancement and denoising module is used to perform real-time, adaptive enhancement and denoising processing when the video image quality deteriorates.
3. A large model-based multi-camera video crowd analysis system according to claim 2, characterized in that: The large model crowd detection and segmentation module specifically includes: The crowd definition module is used to define the crowd detection segmentation parameter as the maximum allowed interval between individuals. Individuals with a distance smaller than this distance are considered to be a crowd. A crowd prompt word template definition module is used to define a crowd prompt word template; A crowd prompt word output module is used to output corresponding prompt words according to the maximum interval allowed between different individuals and the crowd prompt word template adaptively; The annotation module is used to annotate the crowd segmentation results in the form of masks or bounding boxes, and use data annotation tools to add significant annotation boxes into the video content.
4. The multi-camera video crowd analysis system based on a large model according to claim 3 is characterized in that: The large model population composition analysis module specifically includes: Component definition module, used to define different components and component sets; A component prompt word template definition module is used to define a component prompt word template according to the requirements of component analysis; The crowd component type output module is used to combine the crowd component type parameters and confidence, and consider the number of individuals, distance and contact, posture and expression, and environmental factors to perform component analysis on the crowd in the frame image and output the component type of each annotation box.
5. The multi-camera video crowd analysis system based on a large model according to claim 4 is characterized in that: The large model crowd behavior analysis module specifically includes: Behavior definition module, used to define behavior types; A behavior prompt word template definition module is used to define a behavior prompt word template according to the demand characteristics of behavior analysis; The crowd behavior output module is used to generate prompt words for crowd recognition, clarity evaluation, and timestamp and location recording based on the input main camera video and adjacent camera videos as well as crowd segmentation and component analysis results.
6. The multi-camera video crowd analysis system based on a large model according to claim 5 is characterized in that: The context fusion module specifically includes: A camera adjustment module is used to automatically select a camera that can most clearly capture the target group according to the position and movement trajectory of the target group in the video identified by the behavior analysis module, and adjust it to the main camera; The location information storage module is used to record the location information of the target population under the main camera in chronological order, including coordinates, areas, etc.
7. The multi-camera video crowd analysis system based on a large model according to claim 1 is characterized in that: The output result visualization module specifically includes: Data visualization module, used to visualize the results of crowd composition and behavior analysis in the form of charts, images, etc.; The report generation module is used to generate detailed analysis reports according to user needs.
8. A multi-camera video crowd analysis method based on a large model, characterized in that: The following steps are involved: S1, preprocessing the input video clips, uniformly encoding the videos and performing video frame sampling to obtain independent frame images; S2. According to the identification of the target population, the initial main camera and other adjacent cameras are set. According to the identification of the target population and the monitoring requirements, the main camera is set, and other cameras adjacent to the main camera are determined as auxiliary cameras; S3. Define crowd detection segmentation parameters and crowd prompt word templates according to the target crowd identification and actual application scenarios; S4, automatically generate crowd segmentation prompt words based on the current scene and recognition target; S5, inputting the crowd segmentation prompt words and the pre-processed main camera video into the pre-trained large model to perform crowd segmentation processing, and outputting the crowd segmentation result; S6. Based on the results of crowd segmentation and identification of the target population, define crowd components and component prompt word templates; S7. Based on the definition of population components, automatically generate component analysis prompt words; S8, inputting the crowd segmentation results and component analysis prompt words into the large model, performing crowd component analysis, and outputting the component analysis results; S9, based on the component analysis results and the target population, define the behavior analysis prompt word template and automatically generate the behavior analysis prompt word; S10, inputting the main camera video, the adjacent camera videos and the behavior analysis prompt words into the large model, analyzing the crowd behavior, and outputting the behavior analysis results; S11, modifying and optimizing the results of behavior analysis based on contextual information; S12. According to the corrected behavior analysis result and whether the recognition purpose is achieved or the video is ended, determine whether to continue to loop the process from step S4 to step S11 until the stop condition is met.
Citation Information
Patent Citations
Large scale crowd video analysis system and method thereof
CN105447458A
Joint video monitoring method and system thereof
CN101924927A
Behavior recognition method and device, equipment and storage medium
CN113111838A
Abnormal behavior crowd detection method and device based on element mining and video analysis
CN113743184A
Image processing method and related equipment
CN115147492A
Cited By
Multi-mode collaborative video sequence segmentation method
CN120932151A
Method and system for generating monitoring video user attention information based on large model
CN121725409A