A student data intelligent collection method and device based on a human-like eye
By using a human-eye-like intelligent student data collection method and multimodal cameras and feature extraction models, the problems of light changes and occlusion in the science museum exhibition area were solved, enabling accurate visitor identification and activity path analysis, and improving the efficiency of exhibition hall management.
Patent Information
- Application Number
- CN202311030823.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-15
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2043-08-15
AI Technical Summary
The existing video capture system in the science museum exhibition area lacks automatic detection and personnel re-identification functions, and changes in lighting and occlusion affect the recognition efficiency and accuracy, making it difficult to meet the needs of modern management.
A human-eye-like intelligent data acquisition method for students is adopted. Video streams of visitors are acquired through multimodal cameras. Appearance features, human posture features, and virtual trajectories are extracted using a feature extraction model. Kalman filtering and network flow algorithms are combined for matching and path analysis to establish the camera link topology of the exhibition hall and generate visitor activity data.
It enables accurate personnel identification and activity path analysis under varying lighting and occlusion conditions, supporting the exhibition hall's management and services for visitors, and providing data support on behavioral patterns and visitor preferences.
Smart Images

Figure CN117037035B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of personnel re-identification technology, and more specifically, to a method and apparatus for intelligent collection of student data based on human-eye-like perception. Background Technology
[0002] Data collection and identification of visitors to science and technology museums are crucial components of exhibition area management. Traditional manual monitoring and management methods are ill-suited to the demands of modern science and technology museum management. To better understand and manage visitor activity information, a video capture system for personnel re-identification needs to be installed within the exhibition area to achieve precise personnel location and data collection.
[0003] However, existing video capture systems installed in science museum exhibition areas often lack automatic detection and people re-identification functions. Furthermore, due to lighting variations in the exhibition area caused by the exhibits, the video stream contains more noise, potentially causing images to lose detail and clarity, thus affecting the results of people detection and re-identification in the exhibition area. In addition, the gathering of people in the exhibition area often causes non-target pedestrians or non-pedestrians to obstruct the view, reducing the efficiency and accuracy of people re-identification and potentially posing adverse consequences for visitor flow management and exhibition area management in science museums. Summary of the Invention
[0004] To address at least one deficiency or improvement need in the existing technology, this invention provides a human-eye-like intelligent data acquisition method and device for students, which can resist changes in lighting and occlusion, and accurately perform personnel re-identification and activity path analysis for visitors in different exhibition areas of a science and technology museum.
[0005] To achieve the above objectives, according to a first aspect of the present invention, a method for intelligent acquisition of student data based on human-eye analogues is provided, comprising:
[0006] The system acquires video streams of visitors captured by cameras deployed in different exhibition areas within the exhibition hall, generates bounding boxes for each visitor through image detection, and forms a set of bounding boxes to be queried.
[0007] The trained feature extraction model is used to extract the appearance features, human posture features and virtual trajectory of each visitor from the video stream;
[0008] Based on the appearance features, human posture features, and virtual trajectory, the detection box is matched with the visitors to be identified to generate personnel trajectory matching information;
[0009] Based on the personnel trajectory matching information and the pre-established exhibition hall camera link topology, activity data of visitors is generated, including the visitors' identity information and activity trajectory data.
[0010] Furthermore, in the above-mentioned intelligent student data collection method, the process of extracting the physical characteristics of visitors includes:
[0011] Global feature extraction of personnel: Preliminary feature extraction is performed on the video stream images of visitors to generate feature vectors; after the feature vectors are converted into a one-dimensional feature sequence, the final global feature vector is generated after calculation by a self-attention mechanism and a fully connected layer.
[0012] Local feature extraction of personnel: Introducing a learnable set of local prototypes This represents a local classifier that assigns pixels from the global feature vector to the i-th local feature vector, extracts foreground local features from the global feature vector using a cross-attention mechanism, and obtains the final local feature vector after passing through a fully connected layer.
[0013] The global feature vector and the local feature vector are concatenated to generate the appearance features of the visitors.
[0014] Furthermore, in the aforementioned intelligent student data collection method, the process of extracting the human posture features of visitors includes:
[0015] Based on the bounding boxes of visitors, predict the key points of the corresponding human posture and generate human posture features: Where (x) i ,y i ,s i ) represents the i-th position (x) of a total of 17 human body keypoints. i ,y i On, s i The confidence level represents the key points of each human body.
[0016] Furthermore, in the above-mentioned intelligent student data collection method, the process of obtaining the virtual trajectory of the visitors is as follows:
[0017] For each visitor's detection frame, a Kalman filter algorithm is used to generate a virtual trajectory. and the predicted state vector value The steps are as follows:
[0018] (1) Virtual trajectory generation: Based on the Kalman filter algorithm, the posterior state prediction value τ, the posterior covariance matrix P, the state transition matrix F, the observation matrix O, and the noise matrix N between the two modes are generated; in each frame t, the last observed target value is set as The observations that re-trigger the association are represented as The virtual trajectory is then represented as:
[0019]
[0020] (2) Virtual trajectory iteration: The posterior state prediction value τ is moved along the virtual trajectory. The prediction and re-update iterations are performed, and the prediction and re-update operations are as follows:
[0021]
[0022]
[0023] This continues until the observed values on the virtual trajectory match the state vector values calibrated by the latest real observed values.
[0024] Furthermore, in the aforementioned intelligent student data collection method, the step of matching the detection box with the visitors to be identified based on appearance features, human posture features, and virtual trajectories to generate personnel trajectory matching information includes:
[0025] Calculate the distance between the feature vector extracted based on the predicted state vector value and the appearance feature to generate the similarity of the personnel appearance features;
[0026] The positions and confidence levels of key human body points in the human posture features are used as vectors and state vectors in the Kalman filter algorithm to calculate the similarity of key human posture information.
[0027] The similarity of the appearance features of the people and the similarity of the key information of human posture are used to construct a distance matrix in the set of virtual trajectories and the set of detection boxes to be queried. According to the distance matrix, the detection boxes in the current frame are matched with the existing virtual trajectories in the previous frame. If the match is successful, it indicates that the target identity of the virtual trajectory and the detection box is the same, and the personnel trajectory matching information corresponding to the visitors in the detection box is generated.
[0028] Furthermore, the aforementioned intelligent student data collection method employs the Ford-Fulkerson algorithm based on network flow to match the virtual trajectory with the detection box. The calculation method is as follows: G t,t-1 =FFA(D t,t-1 );
[0029] Among them, D t,t-1 G represents the distance matrix; t,t-1 For virtual trajectory set U t-1 and the set of detection boxes to be queried Ω t The best matching; if G is solved t,t-1 A value of 1 indicates a virtual trajectory. Detection boxes in the detection box set The target identities are the same.
[0030] Furthermore, in the aforementioned intelligent student data collection method, the method for establishing the exhibition hall camera link topology is as follows:
[0031] (1) Establish the link topology between different cameras, represented as: G DCMC =(V DCMC E DCMC ), where V DCMC Indicates a camera, V DCMC ={d i |1≤i≤N DCMC};E DCMC E represents the transfer distribution between different cameras. DCMC ={p i,j (Δt)|1≤i≤N DCMC ,1≤j≤N DCMC ,i≠j};
[0032] In the formula, N DCMC d represents the total number of cameras in the link topology. i Let d represent the i-th camera. j Let p represent the j-th camera. i,j (Δt) represents d i and d j The transfer distribution between;
[0033] (2) Establish the topological structure between different exhibition areas, represented as: G EA =(V EA E EA ),V EA ={d i(k) |1≤i≤N DCMC, 1≤k≤A i},
[0034] Among them, A i d represents the number of exhibition areas covered by the i-th camera. i(k) This indicates the k-th exhibition area where the i-th camera is located; It is d i(k) and d j(k) A transition distribution between;
[0035] (3) Perform iterative optimization of the topology, including two steps: updating the time window and relocating the missing personnel;
[0036] Update time window T: For any two exhibition areas, a transition distribution p(Δt) ~ N(μ, α2) is used to adjust the lower and upper time limits of the transition distribution p(Δt) using the prior topology, as follows:
[0037]
[0038] Where μ is a constant, T minAs the lower limit, T max As the upper limit, the time boundary is calculated, and the time window T is updated as follows:
[0039]
[0040] Where α(p(Δt)) is the Gaussian fitting error rate;
[0041] Relocating Disappearing Persons: Based on the topology, a visitor who disappears in the exit area at time t is expected to reappear in the entrance area of another camera at time (t+T). Search for the corresponding person in the topology. If the time slot centers are close to (t+T), the similarity of the appearance features and / or the similarity of the key information of human posture is greater than the preset value. This is considered a reliable correspondence, and the transfer distribution is updated.
[0042] Repeat the above steps until the transfer distribution converges, and execute the above process in all expansion intervals of the link topology until the topology no longer changes or the amount of change in a number of iterations is within a preset range, then stop the iteration process.
[0043] Furthermore, the aforementioned intelligent student data collection method, in which the activity data of visitors is generated based on personnel trajectory matching information and a pre-established exhibition hall camera link topology, includes:
[0044] Based on personnel trajectory matching information, path association is performed between different cameras. The paths to be associated are defined as follows: The virtual trajectories and identity IDs of target visitors are added to the search database. And execute the following process:
[0045] Based on the virtual trajectory of visitors A candidate path library is generated based on the prior exhibition hall camera link topology.
[0046] From the candidate path library Choose the path with the highest similarity With search library The paths in the data are linked to generate visitor activity path information, including timestamps, IDs, and exhibition area numbers.
[0047] According to a second aspect of the present invention, a human-eye-like intelligent student data acquisition device is also provided, comprising:
[0048] The data acquisition unit includes several cameras deployed in different exhibition areas inside the exhibition hall, which are used to collect video streams from visitors.
[0049] The target detection unit is configured to perform image detection on the video stream of visitors captured by the camera, generate a detection box for each visitor, and form a set of detection boxes to be queried.
[0050] A feature extraction unit is configured to extract the appearance features, human posture features and virtual trajectory of each visitor from the video stream using a trained feature extraction model.
[0051] The target matching unit is configured to match the detection box with the visitors to be identified based on the appearance features, human posture features and virtual trajectory, and generate personnel trajectory matching information;
[0052] The activity path generation unit is configured to generate visitor activity data, including visitor identity information and activity trajectory data, based on the personnel trajectory matching information and the pre-established exhibition hall camera link topology.
[0053] Furthermore, the aforementioned intelligent student data acquisition device also includes a control module and a light intensity sensing module in its data acquisition unit;
[0054] The control module is used to control the operation of the camera and adaptively adjust the camera's operating parameters according to the environmental parameters of the exhibition area.
[0055] The light intensity sensing module is used to collect the light intensity in the exhibition area and switch the camera to work in visible light mode or infrared mode according to the light intensity.
[0056] Furthermore, the aforementioned intelligent student data collection device also includes a user interface unit, which is used to visually display visitor activity path data, provide query operation buttons, and setting options.
[0057] According to a third aspect of the present invention, a computer device is also provided, comprising at least one processing unit and at least one storage unit, wherein the storage unit stores a computer program that, when executed by the processing unit, causes the processing unit to perform the steps of any of the above-described intelligent student data collection methods.
[0058] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects:
[0059] The intelligent student data collection method and device based on human-eye-like technology provided by this invention can effectively collect data, identify identities, and generate activity trajectories for exhibition hall visitors. At the same time, it can overcome problems such as changes in lighting and occlusion, which helps exhibition halls manage and serve visitors, understand visitors' behavior patterns and visiting preferences in the exhibition area, and provide data support for comprehensive quality evaluation of personnel in the exhibition hall setting. Attached Figure Description
[0060] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0061] Figure 1 A flowchart illustrating a human-eye-like intelligent student data acquisition method provided in this application embodiment;
[0062] Figure 2 A flowchart illustrating the implementation of an intelligent student data acquisition method based on human-eye analogue provided in this application embodiment;
[0063] Figure 3 A schematic diagram of the tripod installation for deploying binocular multimodal cameras in the exhibition hall;
[0064] Figure 4 A schematic diagram of a high-mounted wall-mounted installation for deploying a binocular multimodal camera in an exhibition hall;
[0065] Figure 5 A schematic diagram of the network structure of the personnel appearance feature extractor provided in the embodiments of this application;
[0066] Figure 6 A logic block diagram of a student data intelligent acquisition device based on human eye-like imaging provided in this application embodiment. Detailed Implementation
[0067] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0068] The terms "first," "second," "third," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0069] Furthermore, to avoid obscuring the understanding of the invention by those skilled in the art, well-known or widely used techniques, elements, structures, and processes may not be described or shown in detail. Although the accompanying drawings illustrate exemplary embodiments of the invention, the drawings are not necessarily drawn to scale, and specific features may be enlarged or omitted to better illustrate and explain the invention.
[0070] Figure 1 This is a flowchart illustrating a human-eye-like intelligent student data acquisition method provided in this embodiment. Figure 2 This is a flowchart illustrating the implementation of a human-eye-like intelligent student data collection method provided in this embodiment. Please refer to it. Figure 1 , 2 The method mainly includes the following steps:
[0071] S1 acquires video streams of visitors collected by cameras deployed in different exhibition areas inside the exhibition hall, generates a detection box for each visitor through image detection, and forms a set of detection boxes to be queried.
[0072] In this method, multiple cameras need to be deployed in different exhibition areas inside the exhibition hall. To ensure the quality of images acquired under different lighting conditions, a binocular multimodal camera (DCMC) is preferred. DCMC can operate in visible light mode and infrared mode respectively, and is represented as: DCMC = {(CVIS1, CIR1); (CVIS2, CIR2); ...; (CVIS...} N CIR N )}, where N represents the number of DCMCs; in different exhibition areas of the exhibition hall, first determine the audience flow routes and main viewing points of the specific exhibition area, and deploy the DCMCs facing the audience flow routes and main viewing points. The installation method is divided into the following categories according to the size of the exhibition area and the architectural structure of the exhibition room: Figure 3 The tripod setup shown is as follows: Figure 4 The high-mounted wall-mounted DCMC is shown. In this embodiment, the tripod-mounted DCMC is deployed in an area less than 150m². 2The support plate in the exhibition area is made of copper-aluminum alloy, which can effectively resist the interference of rapidly changing magnetic fields in the exhibition space. The support plate is 1500mm above the ground, and a pair of DCMCs are fixed parallel to the support plate with a spacing of 150mm. This ensures that the parallax of the two cameras determines the distance and depth of objects, achieving a binocular stereoscopic imaging effect. High-position wall-mounted DCMCs are deployed in areas larger than 150m². 2 In the exhibition area, the mounting pads between the DCMC and the wall are made of copper-aluminum alloy; a pair of DCMCs are fixed to the wall at a height of 2500mm from the ground and a spacing of 150mm, achieving a binocular stereo imaging effect in a large field of view.
[0073] The DCMC (Digital Multi-Modal Camera) uses a binocular multi-modal camera to capture video streams of visitors in different exhibition areas. The DCMC's light intensity sensing module collects the light intensity data value in the current exhibition area and compares it with a pre-set light threshold. When the light threshold is greater than 40% of the light intensity, it indicates that the current exhibition area is well-lit, and the CVIS (Visible Light Sensor) in the DCMC is activated to start capturing the video stream. If the light threshold is less than 40% of the light intensity, it indicates that the current exhibition area is poorly lit, and the CIR (Infrared Sensor) in the DCMC is activated to start capturing the video stream.
[0074] After acquiring several synchronized video streams using the binocular multimodal camera DCMC, a multimodal target detector is used to perform image detection on the image sequence of the video streams. The returned value is the bounding box position information of the visitors, represented as: e m = {x,y,h,w},m∈{visible,infrared}, where (x,y) represents the pixel coordinates of the top-left corner of the detection box, (h,w) represents the width and height of the detection box, and m represents the image modality in which the current detection box is located: visible light mode or infrared mode.
[0075] S2 uses a trained feature extraction model to extract the appearance features, human posture features, and virtual trajectory of each visitor from the video stream;
[0076] (1) Appearance feature extraction
[0077] The process of extracting the physical features of visitors includes global feature extraction and local feature extraction;
[0078] Global feature extraction of personnel: Preliminary feature extraction is performed on the video stream images of visitors to generate feature vectors; after the feature vectors are converted into a one-dimensional feature sequence, the final global feature vector is generated after calculation by a self-attention mechanism and a fully connected layer.
[0079] Local feature extraction of personnel: Introducing a learnable set of local prototypes This represents a local classifier that assigns pixels from the global feature vector to the i-th local feature vector, extracts foreground local features from the global feature vector using a cross-attention mechanism, and obtains the final local feature vector after passing through a fully connected layer.
[0080] Then, the global feature vector and the local feature vector are concatenated to generate the appearance features of the visitors.
[0081] In this embodiment, a personnel appearance feature extractor is used to extract the appearance features of visitors. When using a multimodal camera for image acquisition, the personnel appearance feature extractor comprises three parts: a multimodal fusion unit, a global human body feature extraction module, and a local human body feature extraction module. The training set used to train the personnel appearance feature extractor includes a set of visible light images. and infrared image set in, and These represent infrared and visible light images of visitors from the same video frame, respectively, with N representing the number of images.
[0082] Figure 5 This is a schematic diagram of the network structure of the personnel appearance feature extractor provided in this embodiment. Please refer to [link / reference]. Figure 5 The backbone network of the multimodal feature fusion unit consists of a cross-connected two-stream CNN network with an embedded modality fusion module. Each CNN stream has a three-layer downsampling structure. The multimodal feature fusion unit performs a fusion operation on the infrared and visible light images to generate a feature vector F. vis and F ir ;
[0083] The multimodal fusion process can be divided into the following steps:
[0084] Step 1-1: In the first stage, the two modal images are passed through a 3×3×3 convolutional layer and a BN+ReLU layer before being fed into a DenseNet layer for preliminary feature extraction, resulting in preliminary features, denoted as X. v ′ is and X i ′ r .
[0085] Steps 1-2, Second Stage X v ′ is and X i ′ r Simultaneously, it enters the embedded modality fusion attention module, utilizing global average pooling, as shown below:
[0086]
[0087] Steps 1-3, the compressed vector V vis V ir Entering the fully connected layer, the fused vector V is calculated using the weights ω and the bias term b. fuse The calculation is expressed as: V fuse =ω[(V vis V ir )]+b, where [·] represents the channel connection operation.
[0088] Steps 1-4, vector V fuse Two channel correction vectors R are generated by entering two fully connected layers respectively. vis and R jr The calculation is as follows: R vis =ω1V fuse +b1, R ir =ω2V fuse +b2, where ω1, ω2 and b1, b2 are the weights and biases of the two fully connected layers.
[0089] Steps 1-5, R vis and R ir The two input modalities are integrated using the channel multiplication method, ultimately yielding the feature vector F after modal feature fusion. vis and F ir , is represented as:
[0090] F vis =σ(R) vis )*X' vis F ir =σ(R) ir )*X′ ir (2)
[0091] Where σ represents the activation function and '*' represents the channel multiplication operation, multimodal image feature fusion can calibrate the channels of each modality and extract image background feature information.
[0092] Global feature extraction of personnel: Vector F after modality fusion vis and F ir It can be represented as a sequence of eigenvectors F m =R n ×w×d For m∈{visible, infrared}, the global feature extraction process is the same for both modalities. Taking the visible light modality as an example, the global feature extraction of personnel consists of the following steps:
[0093] Step 2-1, Feature Vector Dimensionality Reduction. The encoder needs a one-dimensional sequence as input, and for a two-dimensional feature sequence F∈R... n ×w×d Performing a flattening and dimensionality reduction operation yields a 1-dimensional feature sequence F of size hw×d. vis =[f 1,vis ;f 2,vis ;...;f hw,vis ], f i,vis ∈R 1×d ;
[0094] Step 2-2, Self-attention mechanism. F vis After passing through the self-attention mechanism, the corresponding Keys, Queries, and Values are generated, as shown in the following mathematical expression:
[0095]
[0096] Where i, j∈1, 2, ..., hw; W Q W K W V These represent the weight matrices for the linear mappings corresponding to Keys, Queries, and Values, respectively; for Its attention weights are based on the dot product similarity between the query and the key, and the calculation is expressed as:
[0097]
[0098] in, These are the corresponding scaling coefficients. After processing with the self-attention mechanism, the resulting feature vector is as follows:
[0099]
[0100] Self-attention weight ζ i,j The mutual dependency information between the i-th pixel and the j-th pixel in the feature input sequence was modeled;
[0101] Steps 2-3, Multi-head Attention Mechanism. The outputs calculated by all self-attention mechanisms are combined to obtain... It contains information about globally neighboring pixels, which is then calculated using a feedforward network, and is represented as:
[0102]
[0103] Here, FFN(·) represents a neural network with two fully connected layers. Residual connections are used before the normalization operation in each layer to mitigate overfitting. After passing through an attention mechanism, the final global feature vector is generated, expressed as:
[0104]
[0105] Steps 2-4: Identity Classification. After the global feature extraction encoder for personnel, an identity classifier ID is added. g Predicting the category distribution yields the probability p. g , is represented as: ID g Cross-entropy loss L between real identity tags cls and triplet loss L tri As the target of classification training;
[0106] Steps 2-5: Encoder training with a total loss function. The trained encoder module extracts global features of people using background context awareness, and uses global average pooling to constrain the global features to satisfy the target loss function, expressed as: f g ∈R 1×d The total loss function L during the training phase of global feature extraction for personnel en Defined as:
[0107]
[0108] Where μ c μ tri To control the hyperparameters of the identity classification loss function and the triad loss function, f represents positive samples with the same identity g , The distance between them, f represents negative samples from different identities g , The distance between them, where α represents the marginal parameter.
[0109] Personnel Local Feature Extraction: This embodiment utilizes a weakly supervised approach for personnel local feature localization and extracts personnel local features based on an encoder. The extraction process consists of the following steps:
[0110] Step 3-1, Cross-Attention Mechanism. First, a set of learnable local paradigms is introduced. It represents a local classifier that judges the feature vector. To determine whether a pixel belongs to the i-th local region, a cross-attention mechanism is used. Foreground local features in the image, obtained after encoder training Cross-attention operation is represented as:
[0111]
[0112] Where i, j∈1, 2, ..., hw; W Q W K W VThese represent the weight matrices for the linear mappings of Keys, Queries, and Values, respectively.
[0113] Step 3-2, generate local features. For each local paradigm Its local mask is calculated as follows:
[0114]
[0115] Where, σ i,j Representing the eigenvector The probability of belonging to the i-th foreground part is calculated, and the attention weights at all hw positions are formed into a mask set represented as: M i =[σ i,1 , σ i,2 , ..., σ i,hw The i-th feature is further obtained through a weighted set, which is defined as the weighted sum of all values, expressed as:
[0116]
[0117] Received The final local features are obtained through two fully connected layers, and are represented as follows:
[0118]
[0119] Step 3-3: Decoder training with total loss function. The decoder for optimizing local feature extraction is trained using local classification loss and triplet loss, with the total loss function L... loc Defined as:
[0120]
[0121] In the process of extracting global and local features of personnel, the personnel appearance feature extractor model is trained by minimizing the overall objective, as follows:
[0122] L cmgl =L En +L loc (13)
[0123] For each image of an unseen identity, after the above steps, the global feature f is... g and local features When connected in series, it can be represented as:
[0124]
[0125] Where [*] represents a chain operation, f cmgl The final appearance features of the person obtained in the feature extraction stage.
[0126] (2) Human posture feature extraction
[0127] For the DCMC deployed within the exhibition hall, = {(CVIS1, CIR1); (CVIS2, CIR2); ...; (CVIS...}, the DCMC is {(CVIS1, CIR1); (CVIS2, CIR2); ...; (CVIS...}. N CIR N Let v i =(CVIS) i CIR i Let represent the video stream acquired by the i-th binocular multimodal camera DCMC. The multimodal target detector in step S1 is used to generate a set of detection boxes to be queried for the current image frame. t represents the timestamp, and i represents the location from the i-th DCMC; this embodiment uses the CmPose algorithm to predict each detection box. The key points of human posture of the visitors are represented by the following set of key human posture information:
[0128]
[0129] Among them, (x i y i s i ) represents the i-th position (x) of a total of 17 human body keypoints. i y i On, s i The confidence level represents the key points of each human body.
[0130] (3) Virtual trajectory generation
[0131] During the personnel tracking process, a set of visitor activity trajectories is established for v i The set of bounding boxes Ω corresponding to each frame of the video stream t A Kalman filter algorithm based on CmaoSORT is used to generate virtual trajectories. and the predicted state vector value The steps are as follows:
[0132] Virtual trajectory generation: Based on the CmaoSORT Kalman filter algorithm, posterior state prediction values τ, posterior covariance matrix P, state transition matrix F, observation matrix O, and noise matrix N between the two modes are generated. In each frame t, the last observed target value is set as... The observations that are retried to the association are represented as The virtual trajectory is represented as:
[0133]
[0134] Virtual trajectory iteration: Following the trajectory, the predicted values based on the posterior τ... The prediction and re-update iterations are performed, and the prediction and re-update operations are as follows:
[0135]
[0136]
[0137] Since the observations on the virtual trajectory match the state vector information calibrated by the latest real observations, the update will no longer be affected by the error accumulated by the Kalman filter algorithm.
[0138] S3 matches the detection box with the visitors to be identified based on the visitors' appearance features, human posture features, and virtual trajectories, generating personnel trajectory matching information;
[0139] In one optional implementation, it includes:
[0140] (1) Calculate the distance between the feature vector extracted based on the predicted state vector value and the appearance feature to generate the similarity of the personnel appearance features;
[0141] In a specific example, cosine distance is used to compute the predicted state vector value. The extracted personnel appearance feature vector and the appearance features extracted in step S2 The feature distance between them is used as the minimum distance as the appearance similarity, and the calculation formula is as follows:
[0142]
[0143] in, Indicates in The system detects the appearance features of target individuals in the image from the trajectory; the smaller the feature distance, the greater the matching similarity.
[0144] (2) The positions and confidence levels of key human body points in the human body posture features are used as vectors and state vectors of the Kalman filter algorithm to calculate the similarity of key human body posture information.
[0145] In a specific example, the human keypoint locations and confidence scores extracted from the human pose features in step S2 are used as the vector and state vector of the Kalman filter algorithm to calculate... and The similarity between them is expressed as:
[0146]
[0147] Where c(·) represents the calculation of the cosine distance between each keypoint, γ p The weight parameter represents the weight of each human body key point, and the κ(·) function is used to determine whether the corresponding human body key point is valid;
[0148] Overall similarity is expressed as:
[0149]
[0150] (3) Construct a distance matrix by combining the similarity of the appearance features of the personnel and the similarity of the key information of human posture in the virtual trajectory set and the set of detection boxes to be queried. Match the detection boxes in the current frame with the existing virtual trajectories in the previous frame according to the distance matrix. If the match is successful, it indicates that the virtual trajectory and the target identity of the detection box are the same, and generate the personnel trajectory matching information corresponding to the visitors in the detection box.
[0151] In one optional implementation, a network flow-based Ford-Fulkerson algorithm is used to match virtual trajectories with bounding boxes. The similarity of human appearance features and key human pose information is then applied to the trajectory set U. t-1 and the set of tests to be queried Ω t Construct the distance matrix D t,t-1 The task of associating candidate detection boxes in the current frame with the existing trajectory set in the previous frame can be viewed as a maximum weight matching problem, and solved using the Ford-Fulkerson algorithm based on network flow. Specifically, each detection box can be considered as a left node, and each existing trajectory as a right node, with the weight being the distance between the two nodes. Then, by solving the maximum weight matching problem, the best matching trajectory between each detection box in the current frame and the previous frame can be obtained. The formula is as follows:
[0152] G t,t-1 =FFA(D t,t-1 ) (twenty two)
[0153] Among them, G t,t-1 For virtual trajectory set U t-1 and the set of detection boxes to be queried Ω t The best matching; if G is solved t,t-1 A value of 1 indicates a virtual trajectory. Detection boxes in the detection box set The target identities are the same.
[0154] S4 generates visitor activity data based on the personnel trajectory matching information and the pre-established exhibition hall camera link topology, including visitor identity information and activity trajectory data.
[0155] In this step, prior matching information is first used to generate visitor trajectory data, and the exhibition hall camera link topology is established, represented as G = (V, E), where y represents the multimodal camera pair (CVIS) in the DCMC. i CIR iE represents the transfer distribution between different camera pairs. The specific establishment method is as follows:
[0156] (1) Establish the link topology between different cameras, including:
[0157] Step 4-1: Initialize the correspondence between associated personnel.
[0158] First, the entire population is divided into multiple subgroups using timestamps, and a series of personnel feature classifiers with overlapping time windows T are trained. Then, other methods (CVIS) are used... i CIR i The personnel feature classifier searches for personnel correspondences. When a visitor is in a certain (CVIS) i CIR i When it disappears from (CVIS) within the time range [tT, t+T], it will be removed from other (CVIS) i CIR i Search for the corresponding relationships for this person. If multiple person feature classifiers overlap with the time range, test all person feature classifiers and select the most reliable one as the optimal correspondence. Use the person's appearance feature similarity score or overall similarity as the criterion for determining the optimal correspondence. i,j >θ sim At that time, θ serves as the most reliable correspondence; sim This represents the preset similarity threshold parameter. In a specific example, θ sim =0.8.
[0159] Step 4-2: Calculate the transfer distribution between DCMCs.
[0160] First, calculate the time difference of correspondences and create a histogram of these time differences. Normalize the histogram using the total number of reliable correspondences. Represent the transition distribution as p(Δt). Histograms of DCMCs with strong connections show significantly dense correspondences within a given time difference, while histograms of DCMCs with weak or no connections show sparse correspondences within the same time difference.
[0161] Step 4-3: Perform a transfer distribution check.
[0162] If two pairs of DCMCs are topologically connected, the transition distribution follows a normal distribution, and the N(μ, σ) of the Gaussian model will be... 2 Fitting the data to the distribution p(Δt), the connection confidence of two pairs of DCMCs is defined as follows:
[0163] conf(p(Δt))=exp(-σ)*(1-α(p(Δt))) (23)
[0164] Where, the confidence range is [0, 1], and α(p(Δt)) represents the model fitting error. When conf(p(Δt)) > θ conf At that time, the two pairs of DCMCs are defined as valid connections.
[0165] Step 4-4: Establish the topology.
[0166] Using the above steps, the link topology between cameras within the exhibition hall can be represented as follows:
[0167]
[0168] Where, N DCMC d represents the total number of DCMCs in the camera link. i Let p represent the i-th DCMC. i,j (Δt) represents d i and d j The distribution of transfers between regions.
[0169] (2) Establish the topological structure between different exhibition areas
[0170] For different exhibition areas, DCMCs (Distributed Management Centers) consider the spatial prior information (entrance / exit areas) between each exhibition area by default during the deployment phase. When two exhibition areas belong to different DCMCs, only the area pairs from exit to entrance are considered. If a visitor goes missing in the exit area at time t, they are likely to appear in the entrance area of different DCMCs within a certain time interval T. Therefore, the correspondence between missing visitors in the entrance areas of different DCMCs is searched within the time range [t, t+T]. Similarly, the visitor feature classifier trained for the entrance area uses reliable correspondences to measure the connectivity confidence of all possible area pairs, and the topological structure between area pairs is described as follows:
[0171]
[0172] Among them, A i d represents the number of exhibition areas covered by the i-th DCMC. i(k) This indicates the k-th exhibition area where the i-th DCMC is located. d i(k) and d j(k) A transition distribution between.
[0173] (3) Iterative optimization of topology
[0174] After obtaining the link topology between different cameras and the topology between different exhibition areas, the personnel re-identification results and network topology are iteratively updated. The specific steps are as follows:
[0175] Step 5-1, update time window T.
[0176] A transition distribution p(Δt) ~ N(μ, σ²) lies between two regions. The lower and upper time bounds of the transition distribution p(Δt) are adjusted using prior topology, as shown in the following equation:
[0177]
[0178] Where μ is a constant, T min As the lower limit, T max As the upper limit, the time boundary is calculated, and the time window T is updated as follows:
[0179]
[0180] Where α(p(Δt)) is the Gaussian fitting error rate. When the fitting error is large, the time window T will become larger.
[0181] Step 5-2: Relocate the missing person.
[0182] Based on the topological structure, a student who disappears in the exit area at time t is expected to reappear in the entrance area of another camera at approximately time (t+T). Using the topological information, we search for the corresponding person in the feature classification model, whose time slot centers are close to (t+T), when the similarity score χ² is high. i,j >θ sim The time is used as a reliable correspondence and the transition distribution is updated.
[0183] Repeat the above steps until the transfer distribution converges, and perform this process for all expansion intervals in the DCMC network topology. If the topology no longer changes or only undergoes minor changes during a few iterations, stop the process.
[0184] Then, by using personnel trajectory matching information and the DCMC network topology, personnel activity paths are correlated to generate visitor activity data. Specifically:
[0185] Based on personnel trajectory matching information, path association is performed between different cameras. The paths to be associated are defined as follows: The virtual trajectories and identity IDs of target visitors are added to the search database. And execute the following process:
[0186] Based on the virtual trajectory of visitors A candidate path library is generated based on the prior exhibition hall camera link topology.
[0187] From the candidate path library Choose the path with the highest similarity With search library The paths in the data are linked to generate visitor activity path information, including timestamps, IDs, and exhibition area numbers.
[0188] It should be noted that although the operations of the methods of the embodiments of this specification are described in a specific order in the above embodiments, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the steps depicted in the flowcharts may be executed in a different order. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0189] This embodiment provides a student data intelligent acquisition device based on human eye-like vision. The device can be implemented in software and / or hardware and can be integrated into electronic devices. Figure 6 This is a logic block diagram of the intelligent student data collection device provided in this embodiment, such as... Figure 6 As shown, the device includes a data acquisition unit, a target detection unit, a feature extraction unit, a target matching unit, an activity path generation unit, and a user interface unit; wherein,
[0190] The data acquisition unit includes several cameras deployed in different exhibition areas within the exhibition hall. These cameras are used to capture video streams from visitors. Preferably, these cameras are binocular multimodal cameras, capable of operating in both visible light and infrared acquisition modes. The camera installation method is described in [link to documentation]. Figure 3 and Figure 4 .
[0191] In a further preferred embodiment, the data acquisition unit also includes a control module and a light intensity sensing module;
[0192] The control module is used to control the operation of the camera and adaptively adjust the camera's operating parameters according to the environmental parameters of the exhibition area;
[0193] The light intensity sensing module is used to collect the light intensity in the exhibition area and switch the camera to work in visible light mode or infrared mode according to the light intensity.
[0194] The target detection unit is configured to perform image detection on the video stream of visitors captured by the camera, generate a detection box for each visitor, and form a set of detection boxes to be queried;
[0195] The feature extraction unit is configured to extract the appearance features, human posture features, and virtual trajectory of each visitor from the video stream using a trained feature extraction model;
[0196] The target matching unit is configured to match the detection box with the visitors to be identified based on the appearance features, human posture features and virtual trajectory, and generate personnel trajectory matching information;
[0197] The activity path generation unit is configured to generate visitor activity data, including visitor identity information and activity trajectory data, based on the personnel trajectory matching information and the pre-established exhibition hall camera link topology.
[0198] The user interface unit is configured to visually display visitor activity path data, provide query operation buttons, and settings options.
[0199] Specific limitations regarding the intelligent student data acquisition device can be found in the limitations of the intelligent student data acquisition method described above, and will not be repeated here. Each module in the aforementioned intelligent student data acquisition device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0200] This embodiment also provides an electronic device, which includes at least one processor and at least one memory. The memory stores a computer program. When the computer program is executed by the processor, it causes the processor to perform the steps of the intelligent student data acquisition method. The specific steps are described above and will not be repeated here. In this embodiment, the types of processor and memory are not specifically limited. For example, the processor can be a microprocessor, a digital information processor, an on-chip programmable logic system, etc.; the memory can be volatile memory, non-volatile memory, or a combination thereof.
[0201] The electronic device can also communicate with one or more external devices (such as a keyboard, pointing terminal, display, etc.), one or more terminals that enable users to interact with the electronic device, and / or any terminal that enables the electronic device to communicate with one or more other computing terminals (such as a network card, modem, etc.). This communication can be performed via an input / output (I / O) interface. Furthermore, the electronic device can also communicate with one or more networks (such as a Local Area Network (LAN), a Wide Area Network (WAN), and / or a public network, such as the Internet) via a network adapter.
[0202] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0203] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0204] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0205] The foregoing description is merely an exemplary embodiment of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Those skilled in the art will readily conceive of embodiments of this disclosure upon considering the specification and practicing the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described herein. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.
[0206] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0207] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A human-eye-like intelligent data acquisition method, characterized in that, include: The system acquires video streams of visitors captured by cameras deployed in different exhibition areas within the exhibition hall, generates bounding boxes for each visitor through image detection, and forms a set of bounding boxes to be queried. The trained feature extraction model is used to extract the appearance features, human posture features, and virtual trajectory of each visitor from the video stream; wherein, the process of obtaining the virtual trajectory is as follows: For each visitor's detection frame, a Kalman filter algorithm is used to generate a virtual trajectory. and the predicted state vector value The steps are as follows: Virtual trajectory generation: Generating posterior state predictions based on the Kalman filter algorithm. Posterior covariance matrix State transition matrix Observation matrix and the noise matrix between the two modes In each frame t In the middle, the target observation value observed last is set as The observations that re-trigger the association are represented as The virtual trajectory is then represented as: Virtual trajectory iteration: predicting the posterior state value Along the virtual trajectory The prediction and re-update iterations are performed, and the prediction and re-update operations are as follows: Until the observed values on the virtual trajectory match the state vector values calibrated by the latest real observed values; Based on the aforementioned appearance features, human posture features, and virtual trajectories, the detection box is matched with the visitors to be identified to generate personnel trajectory matching information; specifically including: Calculate the distance between the feature vector extracted based on the predicted state vector value and the appearance feature to generate the similarity of the personnel appearance features; The positions and confidence levels of key human body points in the human posture features are used as vectors and state vectors in the Kalman filter algorithm to calculate the similarity of key human posture information. The similarity of the appearance features of the people and the similarity of the key information of human posture are used to construct a distance matrix in the set of virtual trajectories and the set of detection boxes to be queried. According to the distance matrix, the detection boxes in the current frame are matched with the existing virtual trajectories in the previous frame. If the match is successful, it indicates that the virtual trajectory and the target identity of the detection box are the same, and the personnel trajectory matching information corresponding to the visitors in the detection box is generated. Based on the personnel trajectory matching information and the pre-established exhibition hall camera link topology, activity data of visitors is generated, including the visitors' identity information and activity trajectory data.
2. The intelligent data acquisition method as described in claim 1, characterized in that, The process of extracting the physical characteristics of the visitors includes: Global feature extraction of personnel: Preliminary feature extraction is performed on the video stream images of visitors to generate feature vectors; after the feature vectors are converted into a one-dimensional feature sequence, the final global feature vector is generated after calculation by a self-attention mechanism and a fully connected layer. Local feature extraction of personnel: Introducing a learnable set of local prototypes , This represents a local classifier that assigns pixels from the global feature vector to the first... i In each local area, the foreground local features in the global feature vector are extracted using a cross-attention mechanism, and the final local feature vector is obtained after passing through a fully connected layer. The global feature vector and the local feature vector are concatenated to generate the appearance features of the visitors.
3. The intelligent data acquisition method as described in claim 1, characterized in that, The process of extracting the human posture features of the visitors includes: Based on the bounding boxes of visitors, predict the key points of the corresponding human posture and generate human posture features, including the position of each key point and its corresponding confidence level.
4. The intelligent data acquisition method as described in claim 1, characterized in that, The Ford-Fulkerson algorithm based on network flow is used to match the virtual trajectory with the detection box. The calculation method is as follows: ; in, Represents the distance matrix; For virtual trajectory sets and the set of detection boxes to be queried The best match; if the solution is found A value of 1 indicates a virtual trajectory. Detection boxes in the detection box set The target identities are the same.
5. The intelligent data acquisition method as described in claim 1, characterized in that, The method for establishing the link topology of the exhibition hall cameras is as follows: (1) Establish the link topology between different cameras, represented as: ,in, Indicates a camera. ; This indicates the transfer distribution between different cameras. ; In the formula, This represents the total number of cameras in the link topology. Indicates the first i One camera, Indicates the first j One camera, express and The transfer distribution between; (2) Establish the topological structure between different exhibition areas, represented as: , , ; in, Indicates the first The number of exhibition areas covered by the cameras. Indicates the first The kth exhibition area where the camera is located; yes and A transition distribution between; (3) Perform iterative optimization of the topology, including two steps: updating the time window and relocating the missing personnel; Update time window T: a transition distribution between any two exhibition areas Adjusting the transfer distribution using prior topology The lower and upper time limits are given by the following formula: in, It is a constant. As the lower limit, As the upper limit, the time boundary is calculated, and the time window T is updated as follows: in, It is the Gaussian fitting error rate; Relocating Missing Persons: Based on the topology, a visitor who disappears at time t in the exit area is expected to reappear at time t in the entrance area of another camera. Appears to the left and right; search for the corresponding relationship of this person in the topology, whose time slot center is close to When the similarity of appearance features and / or key information of human posture between the two is greater than a preset value, it is considered a reliable correspondence and the transition distribution is updated. Repeat the above steps until the transfer distribution converges, and execute the above process in all expansion intervals of the link topology until the topology no longer changes or the amount of change in a number of iterations is within a preset range, then stop the iteration process.
6. The intelligent data acquisition method as described in claim 5, characterized in that, The process of generating visitor activity data based on personnel trajectory matching information and a pre-established exhibition hall camera link topology includes: Based on personnel trajectory matching information, path association is performed between different cameras. The paths to be associated are defined as follows: The virtual trajectories and identity IDs of target visitors are added to the search database. And perform the following process: Based on the virtual trajectory of visitors A candidate path library is generated based on the prior exhibition hall camera link topology. ; From the candidate path library Choose the path with the highest similarity With search library The paths in the data are linked to generate visitor activity path information, including timestamps, IDs, and exhibition area numbers.
7. A data intelligent acquisition device based on human eye-like imaging, characterized in that, include: The data acquisition unit includes several cameras deployed in different exhibition areas inside the exhibition hall, which are used to collect video streams from visitors. The target detection unit is configured to perform image detection on the video stream of visitors captured by the camera, generate a detection box for each visitor, and form a set of detection boxes to be queried. A feature extraction unit is configured to extract the appearance features, human posture features, and virtual trajectory of each visitor from the video stream using a trained feature extraction model; wherein the extraction process of the virtual trajectory is as follows: For each visitor's detection frame, a Kalman filter algorithm is used to generate a virtual trajectory. and the predicted state vector value The steps are as follows: Virtual trajectory generation: Generating posterior state predictions based on the Kalman filter algorithm. Posterior covariance matrix State transition matrix Observation matrix and the noise matrix between the two modes In each frame t In the middle, the target observation value observed last is set as The observations that re-trigger the association are represented as The virtual trajectory is then represented as: Virtual trajectory iteration: predicting the posterior state value Along the virtual trajectory The prediction and re-update iterations are performed, and the prediction and re-update operations are as follows: Until the observed values on the virtual trajectory match the state vector values calibrated by the latest real observed values; A target matching unit is configured to match the detection box with the visitors to be identified based on the appearance features, human posture features, and virtual trajectory, generating personnel trajectory matching information; specifically including: Calculate the distance between the feature vector extracted based on the predicted state vector value and the appearance feature to generate the similarity of the personnel appearance features; The positions and confidence levels of key human body points in the human posture features are used as vectors and state vectors in the Kalman filter algorithm to calculate the similarity of key human posture information. The similarity of the appearance features of the people and the similarity of the key information of human posture are used to construct a distance matrix in the set of virtual trajectories and the set of detection boxes to be queried. According to the distance matrix, the detection boxes in the current frame are matched with the existing virtual trajectories in the previous frame. If the match is successful, it indicates that the virtual trajectory and the target identity of the detection box are the same, and the personnel trajectory matching information corresponding to the visitors in the detection box is generated. The activity path generation unit is configured to generate visitor activity data, including visitor identity information and activity trajectory data, based on the personnel trajectory matching information and the pre-established exhibition hall camera link topology.
8. The intelligent data acquisition device as described in claim 7, characterized in that, The data acquisition unit also includes a control module and a light intensity sensing module; The control module is used to control the operation of the camera and adaptively adjust the camera's operating parameters according to the environmental parameters of the exhibition area. The light intensity sensing module is used to collect the light intensity in the exhibition area and switch the camera to work in visible light mode or infrared mode according to the light intensity.
Citation Information
Patent Citations
Pedestrian recognition system, recognition method and computer readable storage medium
CN109934176A
Method of Tracking Objects in a Video Sequence
US20080181453A1