Identity feature generation method, device, and storage medium
By collecting and fusing behavioral, human, and facial features from various types of cameras, identity features are generated, solving the problem of inaccurate collection caused by mask obstruction and achieving seamless and convenient identity feature collection and recognition.
Patent Information
- Application Number
- CN202011546741.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-24
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2040-12-24
AI Technical Summary
Existing technologies often result in inaccurate identification of individuals in public places due to them wearing masks or covering their faces. Furthermore, the requirement for cameras to remain in place can cause crowds to gather, impacting the convenience of travel.
By capturing video from various types of cameras, behavioral, human, and facial features are extracted, and these features are merged and integrated in a world coordinate system to generate identity features, thus preventing objects from lingering.
It enables seamless identity feature collection, improving the accuracy of identity features and the convenience of travel. It is suitable for scenarios where masks are worn or social distancing is restricted, and ensures the stability and discriminative power of the features.
Smart Images

Figure CN114677608B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a method, device and storage medium for generating identity features. Background Technology
[0002] Nowadays, most public places are equipped with surveillance equipment. This equipment can be used to collect the identity characteristics of individuals, and then apply these characteristics to different scenarios, such as identity authentication or data analysis. Identity characteristics are features that represent the identity of an individual.
[0003] Typically, identity features are the biological characteristics of an object. Taking facial features as an example, during data collection, a camera is usually used to photograph the object, the facial image is extracted from the captured image, and then facial recognition algorithms are used to extract facial features from the facial image.
[0004] For example, in public places during certain periods, people often wear masks for protection. This prevents cameras from capturing complete facial images, making it impossible to extract accurate facial features. Additionally, filming requires the subject to remain in front of the camera for a period of time, which may cause subsequent crowds to accumulate, hindering people's rapid movement and impacting the convenience of travel. Summary of the Invention
[0005] The main objective of this invention is to provide an identity feature generation method, device, and storage medium, which aims to generate identity features seamlessly and improve the accuracy of identity features.
[0006] To achieve the above objectives, embodiments of the present invention provide an identity feature generation method, the method comprising the following steps: acquiring n videos captured by n cameras, wherein the videos contain one or more objects, the n cameras belong to m types, and the installation positions of the cameras of different types are at different distances from the objects being filmed, where n≥m≥2;
[0007] Feature extraction is performed on the n videos in different dimensions to obtain the behavioral features, human body features, and facial features of all objects in each video;
[0008] Based on the installation locations of the n cameras and the n videos, feature merging is performed to determine the behavioral features, human body features, and facial features belonging to the same object;
[0009] The behavioral characteristics, human body characteristics, and facial characteristics belonging to the same object are fused into the identity characteristics of the object.
[0010] In one embodiment, after the behavior features, body features and face features belonging to the same object are fused into the identity feature of the object, the method further comprises:
[0011] acquiring the pre-stored reference identity features in a database;
[0012] calculating the Euclidean distance between each reference identity feature and the identity feature of the object;
[0013] when there is a reference identity feature whose Euclidean distance with the identity feature of the object is less than a first predetermined threshold, determining that the identity authentication of the object is passed, or identifying the object as a known identity person corresponding to the reference identity feature.
[0014] In one embodiment, after the behavior features, body features and face features belonging to the same object are fused into the identity feature of the object, the method further comprises:
[0015] when the public place includes multiple regions and each region is installed with the n cameras, calculating the Euclidean distance between the identity features of all groups of objects in each adjacent two regions, wherein two objects in each group of objects come from different regions;
[0016] when there is a group of objects whose identity features have a Euclidean distance less than a second predetermined threshold in the adjacent two regions, merging the motion trajectories of the group of objects into the motion trajectory of the same object.
[0017] In one embodiment, when the m cameras include a far-end camera and a middle-end camera, and the distance between the installation position of the far-end camera and the object is greater than the distance between the installation position of the middle-end camera and the object, the different dimension feature extraction on the n videos to obtain the behavior features of all objects, the body features of all objects and the face features of all objects in each video comprises:
[0018] performing different dimension feature extraction on the far-end videos in the n videos captured by the far-end camera to obtain the behavior features and body features of each object in each far-end video;
[0019] performing different dimension feature extraction on the middle-end videos in the n videos captured by the middle-end camera to obtain the body features and face features of each object in each middle-end video.
[0020] In one embodiment, the feature merging according to the installation positions of the n cameras and the n videos to determine the behavior features, body features and face features belonging to the same object comprises:
[0021] acquiring a first relative position of the far-end camera and the middle-end camera in a world coordinate system;
[0022] for an i-th object appearing in the far-end video, predicting a first time interval and a first pixel position interval of the i-th object appearing in a far-end video to be shot by the far-end camera in the future according to a motion track of the i-th object in the far-end video, calculating a first real position interval of the i-th object in the world coordinate system according to the first pixel position interval and the first relative position, i≥1;
[0023] acquiring a second time interval and a second pixel position interval of each object appearing in the middle-end video, and calculating a second real position interval of each object in the world coordinate system according to the second pixel position interval and the first relative position;
[0024] if a second time interval corresponding to a j-th object in the middle-end video matches a first time interval corresponding to the i-th object, and a second real position interval corresponding to the j-th object matches a first real position interval corresponding to the i-th object, then merging behavior features and body features of the i-th object and body features and face features of the j-th object into behavior features, body features and face features of the same object, j≥1.
[0025] In an embodiment, when the n cameras further include a near-end camera, and a distance between a mounting position of the middle-end camera and the object is greater than a distance between a mounting position of the near-end camera and the object, the feature extraction of different dimensions on the n videos to obtain behavior features of all objects, body features of all objects and face features of all objects in each video further includes:
[0026] performing feature extraction on a near-end video obtained by the near-end camera in the n videos to obtain face features of each object in each near-end video.
[0027] In an embodiment, the feature merging according to the mounting positions of the n cameras and the n videos to determine behavior features, body features and face features of the same object further includes:
[0028] acquiring a second relative position of the middle-end camera and the near-end camera in the world coordinate system;
[0029] For a u-th object appearing in the mid-end video, a third time interval and a third pixel position interval of the u-th object appearing in a mid-end video to be shot by the mid-end camera in the future are predicted according to a motion track of the u-th object in the mid-end video, a third real position interval of the u-th object in the world coordinate system is calculated according to the third pixel position interval and the second relative position, u≥1;
[0030] A fourth time interval and a fourth pixel position interval of each object appearing in the near-end video are obtained, and a fourth real position interval of each object in the world coordinate system is calculated according to the fourth pixel position interval and the second relative position;
[0031] If a fourth time interval corresponding to a v-th object in the near-end video matches a third time interval corresponding to the u-th object, and a fourth real position interval corresponding to the v-th object matches a third real position interval corresponding to the u-th object, then human body features and face features of the u-th object and face features of the v-th object are merged into human body features and face features belonging to the same object, v≥1.
[0032] In one embodiment, the merging of the behavior features, the human body features and the face features belonging to the same object into the identity features of the object comprises:
[0033] If the object has at least two behavior features, the at least two behavior features are average-pooled, or if the object has at least two human body features, the at least two human body features are average-pooled, or if the object has at least two face features, the at least two face features are average-pooled;
[0034] The obtained behavior features, human body features and face features are stacked to obtain the identity features of the object.
[0035] To achieve the above object, an embodiment of the present application further provides an identity feature generation device, which comprises a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for realizing connection communication between the processor and the memory, and the program is executed by the processor to realize the steps of the foregoing method.
[0036] To achieve the above object, the present application provides a storage medium for computer readable storage, the storage medium storing one or more programs, and the one or more programs are executable by one or more processors to realize the steps of the foregoing method.
[0037] The identity feature generation method, device and storage medium provided by the present application can capture videos of the object by different types of cameras during the walking process of the object, extract behavior features, body features and face features of all objects in each video from the videos, determine the behavior features, body features and face features belonging to the same object, and fuse the behavior features, body features and face features belonging to the same object into the identity features of the object. In this way, the identity features of the object can be collected without awareness during the walking process of the object, and the object can also avoid staying in front of the camera, so that subsequent crowd accumulation will not be caused, thereby improving the convenience of travel.
[0038] In addition, by collecting behavior features, body features and face features in different scales, the richness of biological information can be ensured, and in the case that the object is interfered by noise such as shielding and blurring, the identity features obtained by multi-modal information fusion of the three features can still maintain good discriminability, thereby ensuring the stability and accuracy of the identity features to the greatest extent, having high application value, and being especially suitable for scenes such as wearing masks, dressing easily or requiring social distance.
[0039] In addition, by predicting the position of the object in the world coordinate system, the behavior features, body features and face features collected across cameras can be merged in the time dimension and the space dimension, so that the homology of the three features can be ensured, that is, the effectiveness of multi-modal information fusion is ensured. BRIEF DESCRIPTION OF DRAWINGS
[0040] Figure 1 is a flowchart of the identity feature generation method provided by the embodiment one of the present application.
[0041] Figure 2 is a layout schematic diagram of the camera in the embodiment of the present application.
[0042] Figure 3 is a specific flowchart in the identity authentication and identity recognition scene in the embodiment of the present application.
[0043] Figure 4 is a specific flowchart in the statistical activity track scene in the embodiment of the present application.
[0044] Figure 5 is a structure block diagram of the identity feature generation system in the embodiment of the present application. DETAILED DESCRIPTION
[0045] It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0046] In the following description, the suffixes such as "module", "part", or "unit" used for an element are used only for convenience of explanation of the present application, and have no particular meaning by itself. Thus, "module", "part", or "unit" can be mixedly used.
[0047] Embodiment One
[0048] As shown in the figure, the present embodiment provides an identity feature generation method, which comprises the following steps: Figure 1
[0049] Step 110, acquiring n videos taken by n cameras, the videos containing one or more than one object, the n cameras belonging to m categories, and the distance between the installation position of the camera of different category and the object being taken being different, n≥m≥2.
[0050] The purpose of the present embodiment is to extract the identity feature of the object appearing in the public place. Before implementing the method, n cameras need to be installed in the public place in advance, wherein the public place can be a company park, a community, a street, a shopping mall, etc. In the present embodiment, according to the distance between the installation position of the camera and the object, the n cameras can be divided into m categories, each category containing at least one camera.
[0051] In one application scenario, when m=2, the n cameras can be divided into two categories of far-end cameras and middle-end cameras, and the distance between the installation position of the far-end camera and the object is greater than the distance between the installation position of the middle-end camera and the object. In the present embodiment, the specific value of the distance is not limited. In one example, the distance between the installation position of the far-end camera and the object is about 20 meters, and the distance between the installation position of the middle-end camera and the object is about 10 meters.
[0052] In another application scenario, when m=3, the n cameras can be divided into three categories of far-end cameras, middle-end cameras, and near-end cameras, and the distance between the installation position of the far-end camera and the object is greater than the distance between the installation position of the middle-end camera and the object, and the distance between the installation position of the middle-end camera and the object is greater than the distance between the installation position of the near-end camera and the object. In the present embodiment, the specific value of the distance is not limited. In one example, the distance between the installation position of the far-end camera and the object is about 20 meters, the distance between the installation position of the middle-end camera and the object is about 10 meters, and the distance between the installation position of the near-end camera and the object is about 5 meters.
[0053] Taking the far-end camera as a far-end wide-angle camera, the middle-end camera as a middle-end wide-angle camera, and the near-end camera as a near-end gate camera as an example, please refer to Figure 2 The layout of the n cameras is shown. The n cameras are high and low in the three distance scales of far, middle and near, and the objects are photographed.
[0054] After the n cameras are installed, each camera starts to photograph and sends the photographed video to the device. The device receives n videos, each of which contains one or more objects.
[0055] In step 120, different dimensional feature extraction is performed on the n videos to obtain the behavior features of all objects, the human body features of all objects and the face features of all objects in each video.
[0056] Since the installation position of the far-end camera is farthest from the object, the far-end video photographed by the far-end camera can only identify the behavior and human body features of the object, etc. On this basis, the device can perform different dimensional feature extraction on the far-end video photographed by the far-end camera in the n videos to obtain the behavior features and human body features of each object in each far-end video. The behavior features can represent the gait information of the object, and the human body features can represent the body shape, height, clothing, etc. of the object.
[0057] When extracting the behavior features and human body features from the far-end video, the device can detect the position information of the object in the far-end video by using a pedestrian detection algorithm, which can be an algorithm such as yolo3 (You Only Look Once, general target detection framework), SSD (Single Shot MultiBox Detector, single step multi-frame detector). The device can cut the pedestrian image from the video frame according to the position information, and extract the behavior features of the object by using a behavior feature extraction algorithm, which can be an algorithm such as TSN (Temporal Segment Networks, long-range time segmentation network), I3D (Inflated 3D ConvNet, double-flow inflated 3D convolution network). The device can also cut the pedestrian image from the video frame according to the position information, and extract the human body features of the object by using a human body feature extraction algorithm, which can be an algorithm such as MGN (Mutiple Granularity Network, multi-granularity information fusion network), PCB (Part-based Convolutional Baseline, uniform block network).
[0058] It should be noted that the behavior features and human body features extracted here can be represented by high-dimensional feature vectors obtained by encoding. For example, the behavior features and human body features can be represented by a 256-dimensional feature vector.
[0059] Since the installation position of the middle-end camera is slightly closer to the object, the middle-end video captured by the middle-end camera can recognize slightly fine features such as the human body and the face of the object. On this basis, the device can perform feature extraction of different dimensions on the middle-end video captured by the middle-end camera in the n videos, to obtain the human body features and the face features of each object in each middle-end video. The face features can represent the face shape, skin color, gender, age, and other information of the object.
[0060] When extracting the human body features and the face features from the middle-end video, the device can detect the position information of the object in the middle-end video by using a pedestrian detection algorithm, which can be an algorithm such as yolo3 or SSD. The device can extract the human body features of the object from the pedestrian image in the video frame according to the position information, by using a human body feature extraction algorithm, which can be an algorithm such as MGN or PCB. The device can also detect the position information of the face of the object in the video frame by using a face detection algorithm, which can be an algorithm such as MTCNN (Multi-taskscascade neural network) or RetinaFace (face detection network with feature pyramid). The device can extract the face features of the object from the face image in the video frame according to the position information, by using a face feature extraction algorithm, which can be an algorithm such as ArcFace (face recognition network based on metric learning) or SphereFace (face recognition network based on metric learning).
[0061] It should be noted that the human body features and the face features extracted here can be represented by high-dimensional feature vectors obtained by encoding. For example, the human body features and the face features can be represented by a 256-dimensional feature vector.
[0062] Since the installation position of the middle-end camera is slightly closer to the object, the middle-end video captured by the middle-end camera can recognize slightly fine features such as the human body and the face of the object. On this basis, the device can perform feature extraction of different dimensions on the middle-end video captured by the middle-end camera in the n videos, to obtain the human body features and the face features of each object in each middle-end video. The face features can represent the face shape, skin color, gender, age, and other information of the object.
[0063] When extracting the human body features and the face features from the middle-end video, the device can detect the position information of the object in the middle-end video by using a pedestrian detection algorithm, which can be an algorithm such as yolo3 or SSD. The device can extract the human body features of the object from the pedestrian image in the video frame according to the position information, by using a human body feature extraction algorithm, which can be an algorithm such as MGN or PCB. The device can also detect the position information of the face of the object in the video frame by using a face detection algorithm, which can be an algorithm such as MTCNN (Multi-taskscascade neural network) or RetinaFace (face detection network with feature pyramid). The device can extract the face features of the object from the face image in the video frame according to the position information, by using a face feature extraction algorithm, which can be an algorithm such as ArcFace (face recognition network based on metric learning) or SphereFace (face recognition network based on metric learning).
[0064] It should be noted that the facial features extracted here can be represented by a high-dimensional feature vector obtained by encoding. For example, the facial features can be represented by a 256-dimensional feature vector.
[0065] In step 130, the behavior features, body features and facial features belonging to the same object are determined according to the installation positions of the n cameras and the n videos.
[0066] In step 120, the device can extract the behavior features and body features of all objects from the far-end video, extract the body features and facial features of all objects from the middle-end video, and extract the facial features of all objects from the near-end video. However, the device does not know which behavior features, which body features and which facial features belong to the same object, so the device needs to determine the behavior features, body features and facial features belonging to the same object according to the conditions of time constraints and space constraints.
[0067] Since the object is mostly walked into the public place from the outside, the far-end camera first captures the object, the middle-end camera secondly captures the object, and the near-end camera lastly captures the object. That is, when the far-end camera captures the object, the middle-end camera has not captured the object, so the device can compare the time interval and the position interval of the object appearing in the future predicted from the far-end video with the time interval and the position interval of the object captured by the middle-end camera, so as to determine the same object. Similarly, when the middle-end camera captures the object, the near-end camera has not captured the object, so the device can compare the time interval and the position interval of the object appearing in the future predicted from the middle-end video with the time interval and the position interval of the object captured by the near-end camera, so as to determine the same object. In this way, the features of the same object under the three cameras can be merged to obtain the identity features of the object.
[0068] When the features of the objects in the far-end video and the middle-end video are merged, the behavior features, body features and facial features belonging to the same object are determined according to the installation positions of the n cameras and the n videos, which can include the following steps.
[0069] In step 131, the first relative position of the far-end camera and the middle-end camera in the world coordinate system is obtained.
[0070] The installation positions of the far-end camera and the middle-end camera are known, and the device can pre-establish a world coordinate system and calculate the first relative position of the far-end camera and the middle-end camera in the world coordinate system according to the installation positions of the far-end camera and the middle-end camera.
[0071] Step 132: For the i-th object appearing in the remote video, predict the first time interval and the first pixel position interval of the i-th object in the remote video captured by the remote camera in the future, based on the motion trajectory of the i-th object in the remote video. Calculate the first real position interval of the i-th object in the world coordinate system based on the first pixel position interval and the first relative position, i≥1.
[0072] Multiple objects may appear in the remote video. This embodiment uses the processing flow of the device for the i-th object as an example. The processing flow of other objects is the same as that of the i-th object, and will not be described in detail in this embodiment.
[0073] The device can determine the motion trajectory of an object from its position across multiple video frames in a remote video feed. Then, it uses extrapolation to predict the first time interval and position interval in future remote video shots taken by the remote camera. This position interval is denoted as the first pixel position interval of the object in the camera coordinate system of the remote camera (i.e., the object's position interval in future video frames). Since the transformation relationship between the camera coordinate system of the remote camera and the world coordinate system is known, the device can convert the first pixel position interval into a first real-world position interval in the world coordinate system (i.e., the object's position interval in a public place) based on this transformation relationship and the first relative position.
[0074] Step 133: Obtain the second time interval and second pixel position interval of each object appearing in the mid-range video, and calculate the second real position interval of each object in the world coordinate system based on the second pixel position interval and the first relative position.
[0075] Since the transformation relationship between the camera coordinate system and the world coordinate system of the mid-range camera is known, the device can convert the second pixel position range into the second real position range in the world coordinate system (i.e., the position range of the object in the public place) based on the transformation relationship and the first relative position.
[0076] Step 134: If the second time interval corresponding to the j-th object in the mid-range video matches the first time interval corresponding to the i-th object, and the second real-world location interval corresponding to the j-th object matches the first real-world location interval corresponding to the i-th object, then the behavioral features and human body features of the i-th object and the human body features and facial features of the j-th object are merged into the behavioral features, human body features and facial features of the same object, where j≥1.
[0077] The device can treat objects falling within the same time interval as the same object, and objects falling within the same real-world location interval as the same object. By taking the intersection of the results obtained from the two sets of constraints, the j-th object can be selected, and the j-th object and the i-th object can be determined to be the same object.
[0078] For example, the device determines that object A is likely to appear at the security checkpoint at 5 o'clock according to the remote video, and determines that object B appears at the security checkpoint at 5 o'clock in the mid-end video, and then determines that object A and object B are the same object.
[0079] When the features of the objects in the mid-end video and the near-end video are merged, the features of the behavior, the human body, and the face belonging to the same object are determined according to the installation positions of the n cameras and the n videos, and the following steps can also be included.
[0080] Step 135: Obtain the second relative position of the mid-end camera and the near-end camera in the world coordinate system.
[0081] The installation positions of the mid-end camera and the near-end camera are known, and the device can pre-establish the world coordinate system, and then calculate the second relative position of the mid-end camera and the near-end camera in the world coordinate system according to the installation positions of the mid-end camera and the near-end camera.
[0082] Step 136: For the u-th object appearing in the mid-end video, according to the motion trajectory of the u-th object in the mid-end video, predict the third time interval and the third pixel position interval of the u-th object appearing in the mid-end video shot by the mid-end camera in the future, calculate the third real position interval of the u-th object in the world coordinate system according to the third pixel position interval and the second relative position, and u≥1.
[0083] Multiple objects can appear in the mid-end video, and this embodiment takes the processing flow of the u-th object handled by the device as an example, and the processing flows of other objects are the same as that of the u-th object, which will not be described herein.
[0084] The device can determine the motion trajectory of the object from the positions of the object in multiple video frames in the mid-end video, and then predict the third time interval and the position interval of the object appearing in the mid-end video shot by the mid-end camera in the future by using the extrapolation method. The position interval is recorded as the third pixel position interval of the object in the camera coordinate system of the mid-end camera (i.e. the position interval of the object in the future video frame). Since the conversion relationship between the camera coordinate system of the mid-end camera and the world coordinate system is known, the device can convert the third pixel position interval into the third real position interval in the world coordinate system (i.e. the position interval of the object in the public place) according to the conversion relationship and the second relative position.
[0085] Step 137: Obtain the fourth time interval and the fourth pixel position interval of each object appearing in the near-end video, and calculate the fourth real position interval of each object in the world coordinate system according to the fourth pixel position interval and the second relative position.
[0086] Since the conversion relationship between the camera coordinate system of the proximal camera and the world coordinate system is known, the device can convert the fourth pixel position interval into a fourth real position interval in the world coordinate system (i.e., the position interval of the object in the public place) according to the conversion relationship and the second relative position.
[0087] In step 138, if the fourth time interval corresponding to the vth object in the proximal video matches the third time interval corresponding to the u object, and the fourth real position interval corresponding to the vth object matches the third real position interval corresponding to the u object, the human body features and the face features of the u object and the face features of the vth object are merged into human body features and face features belonging to the same object, v≥1.
[0088] The device can regard the objects falling into the same time interval as the same object, and regard the objects falling into the same real position interval as the same object, and then take the intersection of the results obtained by the two groups of constraint conditions, so as to filter out the vth object and determine that the vth object and the u object are the same object.
[0089] For example, the device predicts that object A is likely to appear at the gate at 5:10 according to the distal video, and determines object B appearing at the gate at 5:10 in the medial video, and then it can be determined that object A and object B are the same object.
[0090] In step 140, the behavior features, the human body features, and the face features belonging to the same object are fused into the identity features of the object.
[0091] If the object has one behavior feature, one human body feature, and one face feature, the device can directly stack the three features to obtain the identity features of the object. For example, each feature is a 256-dimensional feature vector, and the identity features obtained after stacking are a 768-dimensional feature vector.
[0092] If the object has at least two behavior features and / or at least two human body features and / or at least two face features, fusing the behavior features, the human body features, and the face features belonging to the same object into the identity features of the object can include: if the object has at least two behavior features, performing average pooling on the at least two behavior features; or, if the object has at least two human body features, performing average pooling on the at least two human body features; or, if the object has at least two face features, performing average pooling on the at least two face features; and stacking the obtained behavior features, human body features, and face features to obtain the identity features of the object.
[0093] After obtaining the identity features of the object, the device can apply the identity features to different application scenarios, which are described below by taking two application scenarios as examples.
[0094] In a first application scenario, the device can use the identity feature to authenticate the identity of the object, or identify which person the object is. As shown in Figure 3 In this application scenario, the method can further include the following steps.
[0095] Step 150, obtaining the pre-stored reference identity features in the database.
[0096] In this step, the device needs to pre-collect videos of each known identity person appearing in the public place, and store the identity features extracted from these videos into the database as the reference identity features. Since these videos are usually collected in a controlled environment, the extracted reference identity features generally have good discriminability.
[0097] Step 160, calculating the Euclidean distance between each reference identity feature and the identity feature of the object.
[0098] In this step, the method for calculating the Euclidean distance is already very mature, and will not be described herein.
[0099] Step 170, when there is a reference identity feature whose Euclidean distance with the identity feature of the object is less than a first predetermined threshold, determining that the identity authentication of the object is passed, or identifying the object as the known identity person corresponding to the reference identity feature.
[0100] If it is determined that the identity authentication of the object is passed, the device can send a signal to the gate or access control to instruct the gate or access control to be opened, so as to facilitate the object to enter. If the object is identified as a known identity person, the device can send a signal to the display to instruct the display to display the personnel information of the known identity person.
[0101] In a second application scenario, the device can use the identity feature to analyze the activity trajectory of the crowd in the public place. For example, in the transfer center of the subway, the activity trajectory of the crowd is counted, and the flow direction of the crowd is analyzed, so as to reasonably allocate personnel evacuation channels and improve the transfer efficiency. For another example, in a large shopping mall, the crowd residence time and the crowd flow direction are analyzed, which has important reference value for reasonably arranging the exhibition area position and arranging the commodity sales area. As shown in Figure 4 In this application scenario, the method can further include the following steps.
[0102] Step 180, when the public place includes multiple regions, and each region is installed with n cameras, calculating the Euclidean distance between the identity features of all groups of objects in each adjacent two regions, wherein each group of objects includes two objects from different regions.
[0103] The device can divide the public place into multiple regions and assign a corresponding ID (identification) number to each region, such as ID1, ID2...IDN. Each region is installed with n cameras. It should be noted that in this application scenario, only the activity trajectory of the crowd needs to be analyzed, and the identity of the specific object does not need to be recognized, so only the far-end camera and the middle-end camera can be installed, and the near-end camera does not need to be installed.
[0104] The device can use the method described above to calculate the identity features of each object in each region, and then take two adjacent regions to calculate the Euclidean distance between the identity features of all groups of objects in the two adjacent regions, wherein two objects in each group of objects come from different regions. For example, the region of ID1 includes object A and object B, and the region of ID2 includes object C and object D, then the Euclidean distance between object A and object C can be calculated, the Euclidean distance between object A and object D can be calculated, the Euclidean distance between object B and object C can be calculated, and the Euclidean distance between object B and object D can be calculated.
[0105] Step 190, when the Euclidean distance between the identity features of a group of objects in two adjacent regions is less than a second predetermined threshold, the motion trajectory of the group of objects is merged into the motion trajectory of the same object.
[0106] If the Euclidean distance between object A and object B is less than the second predetermined threshold, the device can determine that object A and object B are the same object, and merge the motion trajectories of the object in the two regions, and save the motion trajectory of the object in the database or display it on the display for the analyst to view. The second predetermined threshold can be equal to or different from the first predetermined threshold.
[0107] As shown in FIG. 10, the present embodiment also provides an identity feature generation system, which includes a data acquisition module 510, a data processing module 520 and a result display module 530. Figure 5
[0108] The data acquisition module 510 includes a video data capture module 511 and a video data transmission module 512. The video data capture module 511 can be the far-end camera, the middle-end camera and the near-end camera mentioned above, which is used to complete the video acquisition in the region at the front end. Among them, the far-end camera is responsible for capturing pedestrian video, the middle-end camera is responsible for capturing human body and face video, and the near-end camera is responsible for capturing face video. Then, these videos are transmitted to the data processing module 520 through the video data transmission module 512.
[0109] The data processing module 520 is the core part of the system, and is mainly responsible for the functions of video stream encoding and decoding, video image analysis, and structured data storage. The video image analysis is composed of nine modules, i.e., a pedestrian detection and track tracking module 521, a face detection and track tracking module 522, a face feature extraction module 523, a human body feature extraction module 524, a behavior feature extraction module 525, a feature time dimension matching module 526, a feature space dimension matching module 527, a feature fusion module 528, and a feature matching module 529. The pedestrian detection and track tracking module 521 is responsible for extracting the position information of the object in the video, and the behavior feature extraction module 525 and the human body feature extraction module 524 extract the behavior feature and the human body feature from the video according to the position information. The face detection and track tracking module 522 is responsible for extracting the position information of the object and the position information of the face in the video, and the human body feature extraction module 524 and the face feature extraction module 523 extract the human body feature and the face feature according to the position information. The feature time dimension matching module 526 merges the behavior feature, the human body feature, and the face feature collected under different cameras into the same object according to the appearance order of the object on the time axis; similarly, the feature space dimension matching module 527 merges the behavior feature, the pedestrian feature, and the face feature collected under different cameras into the same object according to the appearance order of the object in the real position space, thereby obtaining three types of biological features belonging to the same object. The feature fusion module 528 is responsible for fusing the three types of biological features to obtain the identity feature of the object. Finally, the feature matching module 529 compares the captured identity feature with the reference identity feature pre-stored in the database to complete the retrieval. The result display module 530 is the terminal part of the system, and includes an interface display module 531 responsible for returning the identity retrieval result of the personnel to the operator through the interface display module 531.
[0110] Embodiment Two
[0111] Embodiment Two of the present application provides an identity feature generation device, which comprises a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing the connection communication between the processor and the memory. When the program is executed by the processor, the specific steps shown in Figure 1 , Figure 3 and Figure 4 are realized.
[0112] Embodiment Three
[0113] Embodiment Three of the present application provides a computer readable storage medium, which stores one or more programs executable by one or more processors to realize the following Figure 1 , Figure 3 and Figure 4The specific steps are shown.
[0114] The identity feature generation method, device and storage medium provided by the embodiment of the present application can capture videos of the object by different types of cameras during the walking process of the object, extract behavior features of all objects, body features of all objects and face features of all objects in each video from the videos, determine the behavior features, body features and face features of the same object, and fuse the behavior features, body features and face features of the same object into the identity features of the object. In this way, the identity features of the object can be collected without awareness during the walking process of the object, and the object can be prevented from stopping in front of the camera, so that subsequent crowd accumulation will not be caused, thereby improving the convenience of travel.
[0115] In addition, by collecting behavior features, body features and face features in different scales, the richness of biological information can be ensured, and in the case that the object is interfered by noise such as shielding and blurring, the identity features obtained by multi-modal information fusion of the three features can still maintain good discriminability, thereby ensuring the stability and accuracy of the identity features to the greatest extent, having high application value, and being especially suitable for scenes such as wearing masks, easy dress or requiring social distance.
[0116] In addition, by predicting the position of the object in the world coordinate system, the behavior features, body features and face features collected across cameras can be merged in the time dimension and the space dimension, so that the homology of the three features can be ensured, that is, the effectiveness of multi-modal information fusion is ensured.
[0117] Those skilled in the art can understand that all or some of the steps in the method disclosed above, the functions of the modules / units in the system and the device can be implemented as software, firmware, hardware and appropriate combinations thereof.
[0118] In hardware implementations, the division of functionality between the functional modules / units referred to in the above description does not necessarily correspond to a division of physical components; for example, one physical component can have multiple functionalities, or one functionality or step can be performed by several physical components in cooperation. Certain physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on computer readable media, which can include computer storage media (or non-transitory media) and communication media (or transitory media). As is well known to those of ordinary skill in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by a computer. Further, it should be appreciated by those skilled in the art that communication media typically embodies computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media.
[0119] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, and are not intended to limit the scope of the present application. Any modification, equivalent replacement and improvement made by those of ordinary skill in the art without departing from the scope and spirit of the present application shall fall within the scope of the present application.
Claims
1. An identity feature generation method, characterized by, The method comprises: acquiring n videos captured by n cameras, the videos containing one or more than one object, the n cameras belonging to m categories, and the distances between the installation positions of the cameras of different categories and the objects being captured being different, n≥m≥2; the n cameras including a far-end camera and a middle-end camera; performing feature extraction of different dimensions on the n videos to obtain behavior features of all objects, body features of all objects, and face features of all objects in each video; performing feature merging according to the installation positions of the n cameras and the n videos to determine behavior features, body features, and face features of the same object; wherein a first time interval at which the same object appears in a far-end video captured by the far-end camera in the future matches a second time interval corresponding to a middle-end video captured by the middle-end camera, and a second real position interval of the same object in a world coordinate system corresponding to the second time interval matches a first real position interval of the same object in a world coordinate system corresponding to the first time interval; fusing the behavior features, the body features, and the face features of the same object into identity features of the object.
2. The method of claim 1, wherein, After the fusing of the behavior features, the body features, and the face features of the same object into the identity features of the object, the method further comprises: acquiring reference identity features pre-stored in a database; calculating Euclidean distances between each reference identity feature and the identity features of the object; when there is a reference identity feature whose Euclidean distance with the identity features of the object is less than a first predetermined threshold, determining that the identity authentication of the object is passed, or identifying the object as a known identity person corresponding to the reference identity feature.
3. The method of claim 1, wherein, After the fusing of the behavior features, the body features, and the face features of the same object into the identity features of the object, the method further comprises: when a public place includes multiple regions, and the n cameras are installed in each region, calculating Euclidean distances between the identity features of all groups of objects in each adjacent two regions, wherein two objects in each group of objects come from different regions; when there is a group of objects whose identity features have a Euclidean distance less than a second predetermined threshold in the adjacent two regions, merging the motion trajectories of the group of objects into the motion trajectory of the same object.
4. The method according to any one of claims 1 to 3, characterized in that, When the distance between the installation position of the far-end camera and the object is greater than the distance between the installation position of the middle-end camera and the object, the performing of the feature extraction of different dimensions on the n videos to obtain the behavior features of all objects, the body features of all objects, and the face features of all objects in each video comprises: performing feature extraction of different dimensions on far-end videos captured by the far-end camera in the n videos to obtain the behavior features and the body features of each object in each far-end video; performing feature extraction of different dimensions on middle-end videos captured by the middle-end camera in the n videos to obtain the body features and the face features of each object in each middle-end video.
5. The method of claim 4, wherein, The behavior features, body features and face features belonging to the same object are determined according to the installation positions of the n cameras and the n videos, and the behavior features, body features and face features belonging to the same object are determined by: obtaining a first relative position of the remote camera and the middle camera in a world coordinate system; for the i-th object appearing in the remote video, predicting a first time interval and a first pixel position interval of the i-th object appearing in a remote video to be shot by the remote camera in the future according to a motion track of the i-th object in the remote video, and calculating a first real position interval of the i-th object in the world coordinate system according to the first pixel position interval and the first relative position, i≥1; obtaining a second time interval and a second pixel position interval of each object appearing in the middle video, and calculating a second real position interval of each object in the world coordinate system according to the second pixel position interval and the first relative position; if the second time interval corresponding to the j-th object in the middle video matches the first time interval corresponding to the i-th object, and the second real position interval corresponding to the j-th object matches the first real position interval corresponding to the i-th object, then the behavior features and body features of the i-th object and the body features and face features of the j-th object are merged into the behavior features, body features and face features belonging to the same object, j≥1.
6. The method of claim 4, wherein, When the n cameras further include a near camera, and the distance between the installation position of the middle camera and the object is greater than the distance between the installation position of the near camera and the object, the feature extraction of different dimensions is performed on the n videos to obtain the behavior features of all objects, the body features of all objects and the face features of all objects in each video, and the feature extraction of different dimensions further includes: performing feature extraction on a near video obtained by the near camera in the n videos to obtain the face features of each object in each near video.
7. The method of claim 6, wherein, The behavior features, body features and face features belonging to the same object are determined according to the installation positions of the n cameras and the n videos, and the behavior features, body features and face features belonging to the same object are determined by: obtaining a second relative position of the middle camera and the near camera in a world coordinate system; for the u-th object appearing in the middle video, predicting a third time interval and a third pixel position interval of the u-th object appearing in a middle video to be shot by the middle camera in the future according to a motion track of the u-th object in the middle video, and calculating a third real position interval of the u-th object in the world coordinate system according to the third pixel position interval and the second relative position, u≥1; obtaining a fourth time interval and a fourth pixel position interval of each object appearing in the near video, and calculating a fourth real position interval of each object in the world coordinate system according to the fourth pixel position interval and the second relative position; If a fourth time interval corresponding to a vth object in the near-end video matches a third time interval corresponding to a u th object, and a fourth real position interval corresponding to the vth object matches a third real position interval corresponding to the u th object, then the human body features and the face features of the u th object and the face features of the vth object are merged into human body features and face features of the same object, v≥1.
8. The method of claim 6, wherein, The fusing of the behavior features, the human body features and the face features of the same object into the identity features of the object comprises: If the object has at least two behavior features, then the at least two behavior features are average-pooled; or, If the object has at least two human body features, then the at least two human body features are average-pooled; or, if the object has at least two face features, then the at least two face features are average-pooled; The obtained behavior features, human body features and face features are stacked to obtain the identity features of the object.
9. An identity feature generation device, comprising a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing connection communication between the processor and the memory, the program being executed by the processor to realize the steps of the identity feature generation method according to any one of claims 1 to 8.
10. A storage medium for computer-readable storage, the storage medium storing one or more programs executable by one or more processors to realize the steps of the identity feature generation method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Pedestrian identification system based on multi-ball multi-gun camera array
CN105979210A
Access control system with integration of face recognition and gait recognition
CN109920111A