Identity recognition method, model training method, device, equipment and storage medium
The multi-attribute classification model addresses the overfitting and generalization issues in deep learning systems by using shared networks and attention mechanisms, facilitating adaptable and efficient identity recognition across various security monitoring scenarios.
Patent Information
- Application Number
- CN202011448967.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-09
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2040-12-09
AI Technical Summary
Existing deep learning identity recognition systems are difficult to obtain sample sets in specific scenarios, the model is prone to overfitting, and the generalization ability is poor, making it difficult to adapt to the needs of multiple monitoring scenarios.
A multi-attribute classification model is adopted to train by obtaining public image data sets, construct a sample set, use a shared backbone network, determine multiple attributes of the target person, and identify whether the identity meets the entry conditions based on the standard attributes of the monitoring scene, simplify the sample set acquisition process and reduce the risk of overfitting.
It improves the generalization ability of the model, can adapt to the needs of different monitoring scenarios, simplifies the sample set acquisition process, reduces the risk of model overfitting, and improves the flexibility and efficiency of identity identification.
Smart Images

Figure CN114612813B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of security monitoring, and particularly to an identity recognition method, a model training method, a device, a device and a storage medium. Background Art
[0002] In recent years, the technology in the field of security monitoring has developed rapidly, and person recognition is a typical application in the field of security monitoring. Some places only allow staff with specific identities and wearing specific clothing to enter, and do not allow unqualified people to enter. If a person who does not meet the clothing requirements appears in this area, an alarm needs to be triggered. For example, in a military jurisdiction, only military personnel wearing designated clothing are allowed. When the system detects a person whose clothing does not meet the requirements, it means that a suspicious person has been detected, and the system needs to alarm and request the staff to verify the identity of the suspicious person. The identity recognition system using traditional image processing methods has low accuracy, so the existing identity recognition systems mainly adopt deep learning methods.
[0003] Currently, most deep learning systems need to collect a large amount of data as a training set in each application scenario and train a model applicable to the specified scenario. However, such a model has the following disadvantages: it is very difficult to obtain the sample set in a specific scenario, the trained model is prone to overfitting, and the generalization ability of the model is poor, making it difficult to meet the monitoring requirements of more monitoring scenarios. Summary of the Invention
[0004] The main purpose of the embodiments of the present application is to propose an identity recognition method, a model training method, a device, a device and a storage medium, aiming to simplify the process of obtaining the sample set, reduce the risk of model overfitting, and improve the generalization ability of the model to meet the monitoring requirements of more monitoring scenarios.
[0005] To achieve the above object, the embodiments of the present application provide an identity recognition method, including: obtaining a video image in a monitoring scenario; if a target person appears in the video image, determining multiple attributes of the target person according to a pre-trained multi-attribute classification model, where the multi-attribute classification model is trained according to a pre-constructed sample set, and the sample set includes a number of images labeled with attributes; determining the standard attributes of the identity that meet the entry conditions of the monitoring scenario; and identifying whether the identity of the target person meets the entry conditions according to the multiple attributes of the target person and the standard attributes.
[0006] To achieve the above object, an embodiment of the present application provides a training method for a multi-attribute classification model, including: obtaining a publicly available image dataset; annotating multiple attributes of the people in the images that meet the preset annotation conditions in the image dataset to construct the sample set; determining the structure of the network and configuring the network hyperparameters of the network; training the network configured with the network hyperparameters according to the sample set to obtain the multi-attribute classification model.
[0007] To achieve the above object, an embodiment of the present application provides a training device for a multi-attribute classification model, including: an obtaining module for obtaining a publicly available image dataset; an annotating module for annotating multiple attributes of the people in the images that meet the preset annotation conditions in the image dataset to construct the sample set; a configuring module for determining the structure of the network and configuring the network hyperparameters of the network; a training module for training the network configured with the network hyperparameters according to the sample set to obtain the multi-attribute classification model.
[0008] To achieve the above object, an embodiment of the present application further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above identity recognition method.
[0009] To achieve the above object, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, and the computer program realizes the above identity recognition method when executed by a processor.
[0010] In the embodiments of the present application, a video image in a monitoring scenario is obtained; if a target person appears in the video image, multiple attributes of the target person are determined according to a pre-trained multi-attribute classification model, where the multi-attribute classification model is trained according to a pre-constructed sample set, and the sample set includes several images annotated with attributes; a standard attribute for an identity that meets the entry condition of the monitoring scenario is determined; according to the multiple attributes of the target person and the standard attribute, it is identified whether the identity of the target person meets the entry condition. That is to say, compared with the models applicable to specific scenarios in the prior art, in this embodiment, the entry condition for an identity that meets the monitoring scenario is defined by the standard attribute, and different monitoring scenarios can define different standard attributes, so that in this embodiment, a multi-attribute classification model can be trained to adapt to the monitoring requirements of different monitoring scenarios, and the same multi-attribute classification model can be applicable to different monitoring scenarios, which is beneficial to improving the generalization ability of the model. Moreover, there is no need to obtain training data in a specific scenario to train an attribute classification model for the specific scenario, which is beneficial to reducing the risk of model overfitting and avoiding obtaining training data in a specific scenario where it is not easy to obtain training data, that is, to a certain extent, simplifying the process of obtaining the sample set. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 is a flowchart of the identity recognition method mentioned in the first embodiment of the present application;
[0012] Figure 2 is a schematic diagram of the multi-task classification model mentioned in the first embodiment of the present application and the single-task classification model in the prior art;
[0013] Figure 3 is a schematic diagram of introducing an attention mechanism into the multi-attribute classification model mentioned in the second embodiment of the present application;
[0014] Figure 4(a) is the original unannotated image mentioned in the second embodiment of the present application;
[0015] Figure 4(b) is a schematic diagram of different regions marked with different colors mentioned in the second embodiment of the present application;
[0016] Figure 5 is a flowchart of the implementation manner of determining multiple attributes of a target person according to a pre-trained multi-attribute classification model mentioned in the second embodiment of the present application;
[0017] Figure 6 is a schematic diagram of the mask image corresponding to the upper body region mentioned in the second embodiment of the present application;
[0018] Figure 7 is a flowchart of the training method of the multi-attribute classification model mentioned in the third embodiment of the present application;
[0019] Figure 8 It is a schematic diagram of a training device for a multi-attribute classification model mentioned in the fourth embodiment of the present application;
[0020] Figure 9 It is a schematic diagram of the structure of an electronic device mentioned in the fifth embodiment of the present application. Specific embodiments
[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the embodiments of the present application will be described in detail below with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in the embodiments of the present application, many technical details are presented for the better understanding of the present application by readers. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented. The following division of each embodiment is for convenience of description and should not constitute any limitation to the specific implementation manner of the present application. Each embodiment can be combined and cross-referenced with each other on the premise of no contradiction.
[0022] In the embodiments of the present application, it is considered that in the related art, most deep learning systems need to collect a large amount of data as a training set in each application scenario and train a model applicable to the specified scenario. However, the inventors of the present application found that such models have the following disadvantages:
[0023] (1) It is very difficult to obtain a high-quality sample set in a specific scenario. Training a deep learning network model requires a large amount of diverse data. Some places are classified places, and the amount of data that can be obtained from these classified places is limited. At the same time, the pattern of the data obtained in a specific scenario is relatively single, with limited diversity, which is not conducive to the training of the deep learning network model and is extremely likely to cause the network model to overfit.
[0024] (2) The model applicable to the specified scenario has a high accuracy when applied to the specified scenario, but if it is switched to other similar scenarios, the model may completely fail. For example, when the model applied to the reference room of Hospital A is migrated to the reference room of Hospital B, the styles and colors of the uniforms of the staff in Hospital B may be different from those in Hospital A. However, since the model only focuses on the characteristics of the uniforms of the staff in Hospital A, when the model is applied to the reference room of Hospital B, the model may completely fail. If the model needs to be applied to Hospital B, data needs to be collected in the reference room of Hospital B and the model needs to be retrained. This limits the large-scale deployment of the model, and the generalization ability of the model is poor.
[0025] To solve the above technical problems that it is very difficult to obtain a sample set in a specific scenario, the trained model is prone to overfitting, and the generalization ability of the model is poor, the embodiments of the present application provide the following identity recognition method, aiming to simplify the process of obtaining the sample set, reduce the risk of model overfitting, and improve the generalization ability of the model.
[0026] The first embodiment of this application relates to an identity recognition method, which is applied to an electronic device; among them, the electronic device can be a server. The application scenarios of this embodiment can include, but are not limited to: scenarios with security monitoring requirements such as hospital data rooms, police station data rooms, bank data rooms, military jurisdiction areas, prisons, and factory production workshops. The implementation details of the identity recognition method in this embodiment will be specifically described below. The following content is only the implementation details provided for easy understanding and is not necessary for implementing this solution.
[0027] The flowchart of the identity recognition method in this embodiment can refer to Figure 1 , including:
[0028] Step 101: Obtain video images in the monitoring scenario.
[0029] Among them, the monitoring scenario can be the above-mentioned hospital data room, police station data room, bank data room, military jurisdiction area, prison, and factory production workshop, etc. Several monitoring cameras can be deployed in the monitoring scenario to collect video images in the monitoring scenario and transmit the collected video images to the server, so that the server can obtain the video images in the monitoring scenario. In a specific implementation, several monitoring cameras can collect video images in the monitoring scenario in real time, so that the server can obtain the video images in the monitoring scenario in real time to improve the reliability of monitoring.
[0030] Step 102: If a target person appears in the video image, determine multiple attributes of the target person according to a pre-trained multi-attribute classification model.
[0031] Among them, the target person can be understood as any person who appears in the video image. That is to say, if any person appears in the video image, it can be determined that a target person appears in the video image. In a specific implementation, it can also be understood that the server performs target detection on the video image, and when the detected target is a person, it is determined that a target person appears in the video image.
[0032] In an example, the method for determining whether a target person appears in the video image can be: using a pre-trained pedestrian detection model to detect whether the target in the video image is a person. The training method of the pedestrian detection model will be described below:
[0033] (1) Establishment of the image dataset: Obtain publicly available image datasets. That is to say, the above-mentioned image datasets can use a large number of publicly available datasets. In actual deployment scenarios, the workload of collecting data is large, and the diversity of data is limited. Using publicly available image datasets can simplify the complex process of making image datasets without having to collect data in actual deployment scenarios, and more data can be used to train the model. However, in specific implementations, images in multiple monitoring scenarios can also be collected to construct an image dataset.
[0034] (2) Training of the pedestrian detection model: Select a target detection network structure, configure network hyperparameters, and use the constructed image dataset to train the pedestrian detection model. Among them, the target detection network structure can be a one-stage target detection network structure or a two-stage target detection network structure. The one-stage target detection network structure can include, but is not limited to, Single Shot Detector (SSD for short), You Only Look Once (YOLO for short), and Fully Convolutional One-Stage Object Detection (FCOS for short). The two-stage target detection network structure can be Faster Region CNN (Faster RCNN for short).
[0035] Optionally, to improve the reliability of the trained pedestrian detection model, after training the pedestrian detection model, it can also include:
[0036] (3) Performance evaluation of the pedestrian detection model: Evaluate the performance of the trained pedestrian detection model. If the performance does not meet the application requirements, the above step (2) can be returned to re-select the target detection network structure or re-configure the network hyperparameters to retrain the pedestrian detection model.
[0037] Optionally, to improve the running efficiency of the trained pedestrian detection model, in step (3), if the performance meets the application requirements, the following steps can also be carried out:
[0038] (4) Quantization and compression of the pedestrian detection model: The data processed by this pedestrian detection model is video data. Due to limited hardware computing power, to ensure the running efficiency of the model, the trained pedestrian detection model can be quantized and compressed. The acceleration and quantization compression of the model can effectively improve the running efficiency of the model.
[0039] In this embodiment, if a target person appears in the video image, multiple attributes of the target person are determined according to a pre-trained multi-attribute classification model. Among them, the multi-attribute classification model is trained according to a pre-constructed sample set, and the sample set includes several images labeled with attributes. The above multi-attribute classification model can be understood as a multi-task classification model. Each classification task can be understood as the classification of one attribute, and multiple classification tasks can be understood as the classification of multiple attributes. Compared with a single-task classification model, multiple classification tasks share the same backbone network. Multiple-task learning can promote the model to learn shared feature representations and improve the generalization ability of the model. The above multiple attributes can include, but are not limited to: whether wearing a hat, whether wearing epaulets, the color of the clothes, the texture of the clothes, the style of the clothes.
[0040] To facilitate understanding the differences between the multi-task classification model in this embodiment and the single-task classification model in the prior art, reference can be made to Figure 2 . Among them, the single-task classification model is the classification model 1, classification model 2... classification model n in the figure. The classification task of classification model 1 is the classification of the attribute of the clothes style, the classification task of classification model 2 is the classification of the attribute of the clothes color, and the classification task of classification model n is the classification of the attribute of whether wearing a hat. The classification tasks of the multi-task classification model are: the classification of multiple attributes such as clothes style, clothes color, and whether wearing a hat. That is to say, in the prior art, each single-task classification model requires a backbone network, and multiple backbone networks are required to complete multi-classification tasks. In this embodiment, however, the multi-task classification model only requires one backbone network, and one backbone network is shared to complete multi-classification tasks, which is beneficial to improving the operation efficiency of the network.
[0041] In an example, the training method of the multi-attribute classification model can be as follows:
[0042] (1) Obtain a publicly available image dataset. Among them, this image dataset can be the image dataset constructed when training the above pedestrian detection model. In specific implementation, the above image dataset can use a large number of publicly available datasets. In the actual deployment scenario, the workload of collecting data is large, and the diversity of data is limited. In this embodiment, a large number of publicly available datasets can be used when training the model, without having to collect data in the actual deployment scenario, which simplifies the complicated process of making the image dataset and can use more data to train the model.
[0043] (2) Annotate multiple attributes of the people in the images that meet the preset annotation conditions in the image dataset to construct a sample set; among them, the preset annotation conditions can be set according to actual needs. For example, they can be: the people in the image are not blocked, the area occupied by the people in the image is larger than the preset area, the number of body parts shown by the people in the image exceeds the preset number, etc. Both the above-mentioned preset area and preset number can be set according to actual needs, and this embodiment does not make specific limitations on this. In specific implementation, the multiple attributes annotated for the people in the image include but are not limited to: the style, color, texture of the clothes worn by the person, whether the person wears a hat, whether the person wears epaulets, etc. That is to say, the attributes of some people in the image dataset can be annotated to construct a sample set of human attributes.
[0044] (3) Determine the structure of the network and configure the network hyperparameters of the network. Among them, the structure of the network includes a backbone network, and the backbone network can select MobileNet. MobileNet belongs to a lightweight network and has high operating efficiency.
[0045] (4) Train the network configured with network hyperparameters according to the sample set to obtain a multi-attribute classification model.
[0046] Optionally, in order to improve the reliability of the trained multi-attribute classification model, after training the multi-attribute classification model, it may further include:
[0047] (5) Evaluate the performance of the trained multi-attribute classification model. If the model performance does not meet the application requirements, redesign the backbone network of the multi-attribute classification model or reconfigure the network hyperparameters, and retrain the multi-attribute classification model.
[0048] Optionally, in order to improve the operating efficiency of the trained multi-attribute classification model, in step (5), if the performance meets the application requirements, the following steps can also be performed:
[0049] (6) Model quantization compression. For example, the trained multi-attribute classification model can be quantized and compressed using TensorRT. The acceleration and quantization compression of the model can effectively improve the operating efficiency of the model. TensorRT is a high-performance deep learning inference optimizer that can provide low-latency and high-throughput deployment inference for deep learning applications. TensorRT can be used to accelerate inference for ultra-large data centers, embedded platforms, or autonomous driving platforms.
[0050] Step 103: Determine the standard attributes of the identities that meet the entry conditions of the monitoring scenario.
[0051] Among them, for the monitoring requirements of different monitoring scenarios, the identities of the people allowed to enter different monitoring scenarios may be different. Therefore, different monitoring scenarios may correspond to different standard attributes.
[0052] In one example, the monitoring scenario is the data room of Hospital A. The identities of the people allowed to enter the data room of Hospital A are doctors, nurses, and hospital logistics staff. Among them, doctors and nurses both wear long white work uniforms, and logistics staff all wear short blue tops and blue trousers. The standard attributes of the identities that meet the entry conditions for the data room of Hospital A include: long white work uniforms (standard attributes of doctors and nurses), short blue tops, and blue trousers (standard attributes of logistics staff).
[0053] In another example, the monitoring scenario is the production workshop of Factory A. The production workshop of the factory is a dangerous area, and non-factory staff are strictly prohibited from entering. The staff in the production workshop of this factory include three types: Type A workers wearing blue tops and grey trousers, Type B workers wearing red tops and red trousers, and Type C workers wearing orange vests and orange trousers. The standard attributes of the identities that meet the entry conditions for the production workshop of Factory A include: blue tops and grey trousers (standard attributes of Type A workers), red tops and red trousers (standard attributes of Type B workers), orange vests and orange trousers (standard attributes of Type C workers).
[0054] In specific implementation, the server can pre-store the standard attributes of the identities that meet the entry conditions for the monitoring scenario. For example, if the monitoring scenario is the data room of Hospital A, the server can be the monitoring server for the data room of Hospital A, and the standard attributes of the identities that meet the entry conditions for the data room of Hospital A can be pre-stored in this monitoring server. Another example is that if the monitoring scenario is the production workshop of Factory A, the server can be the monitoring server for the production workshop of Factory A, and the standard attributes of the identities that meet the entry conditions for the production workshop of Factory A can be pre-stored in this monitoring server.
[0055] Step 104: Identify whether the identity of the target person meets the entry conditions based on the multiple attributes of the target person and the standard attributes.
[0056] Specifically, the server can match the multiple attributes of the target person with the standard attributes. If the match is successful, it is identified that the identity of the target person meets the entry conditions; otherwise, it is identified that the identity of the target person does not meet the entry conditions. Among them, the matching method can be: the server compares the multiple attributes of the target person with the standard attributes. If there are attributes in the multiple attributes of the target person that are the same as the standard attributes, it can be considered that the identity of the target person meets the entry conditions.
[0057] In one example, the standard attributes of identities that meet the entry conditions for the monitoring scenario include multiple standard attributes corresponding to multiple identities. Based on the multiple attributes of the target person and the standard attributes, the method for identifying whether the identity of the target person meets the entry conditions can be as follows: The server matches the multiple attributes of the target person with each standard attribute respectively. If the multiple attributes of the target person match successfully with the standard attributes corresponding to any one identity, it is identified that the identity of the target person meets the entry conditions. That is to say, the server matches the multiple attributes of the target person with each standard attribute in sequence until a successful match is determined to indicate that the identity of the target person meets the entry conditions, or until a failed match is determined to indicate that the identity of the target person does not meet the entry conditions.
[0058] For example, the standard attributes of identities that meet the entry conditions for the monitoring scenario set in the production workshop of Factory A mentioned in the above example include: the standard attributes of Work Type A, the standard attributes of Work Type B, and the standard attributes of Work Type C. That is, the standard attributes of identities that meet the entry conditions for the monitoring scenario include 3 standard attributes corresponding to 3 identities. The server can first match the multiple attributes of the target person with the standard attributes of Work Type A, that is, determine whether there are attributes in the multiple attributes of the target person that are the same as the standard attributes of Work Type A. If there are, it is considered that the multiple attributes of the target person match successfully with the standard attributes of Work Type A. If there are no attributes in the multiple attributes of the target person that are the same as the standard attributes of Work Type A, the multiple attributes of the target person can be further matched with the standard attributes of Work Type B, that is, determine whether there are attributes in the multiple attributes of the target person that are the same as the standard attributes of Work Type B. If there are, it is considered that the multiple attributes of the target person match successfully with the standard attributes of Work Type B. If there are no attributes in the multiple attributes of the target person that are the same as the standard attributes of Work Type B, the multiple attributes of the target person can be further matched with the standard attributes of Work Type C, that is, determine whether there are attributes in the multiple attributes of the target person that are the same as the standard attributes of Work Type C. If there are, it is considered that the multiple attributes of the target person match successfully with the standard attributes of Work Type C. If not, it means that the multiple attributes of the target person do not match any of the above 3 standard attributes, and it can be identified that the identity of the target person does not meet the entry conditions.
[0059] In one example, the way to match multiple attributes of a target person with each standard attribute can be as follows: Determine the priorities of multiple standard attributes, and in accordance with the priorities of the multiple standard attributes, sequentially match the multiple attributes of the target person with each standard attribute. Among them, the priorities of the multiple standard attributes can be preset according to actual needs and stored in the server. For example, the priorities of the standard attributes of the above-mentioned job type A, job type B, and job type C from high to low are in turn: the standard attribute of job type C, the standard attribute of job type B, and the standard attribute of job type A. Then when the server performs the matching, it can first match the multiple attributes of the target person with the standard attribute of job type C. If the matching is unsuccessful, then match the multiple attributes of the target person with the standard attribute of job type B. If the matching is still unsuccessful, then match the multiple attributes of the target person with the standard attribute of job type A. By setting priorities for multiple standard attributes, it is beneficial to match the multiple attributes of the target person with each standard attribute in a reasonable order.
[0060] In one example, the priority can be determined based on the actual number of people corresponding to multiple identities in the monitoring scenario; among them, the higher the actual number of people corresponding to an identity, the higher the priority of the standard attribute corresponding to it. For example, the actual number of people corresponding to the above-mentioned job type A is 50, the actual number of people corresponding to job type B is 60, and the actual number of people corresponding to job type C is 70. That is to say, in the production workshop of factory a above, theoretically there are 50 workers belonging to job type A, 60 workers belonging to job type B, and 60 workers belonging to job type C. Then the priorities of the three standard attributes corresponding to the above three job types from high to low are in turn: the standard attribute of job type C, the standard attribute of job type B, and the standard attribute of job type A. Since the number of workers belonging to job type C among the workers in the production workshop of factory a is the largest, the probability that a worker entering the production workshop of factory a belongs to job type C is relatively high. Therefore, when performing the matching, it is easier to match successfully by preferentially matching the multiple attributes of the target person with the standard attribute with a higher priority, and thus there is no need to perform the matching of the standard attribute of the next priority, which is beneficial to improving the speed of identity recognition.
[0061] In the specific implementation, if it is recognized that the identity of the target person does not meet the entry conditions, an alarm mechanism can be triggered to remind relevant personnel that there may be illegal personnel intrusion in the monitoring scenario, so as to conduct verification in a timely manner. Among them, the alarm mechanism can be set according to actual needs, and this embodiment does not make specific limitations on this.
[0062] For the convenience of understanding the embodiments, the following will be described with two specific monitoring scenarios:
[0063] Monitoring scenario 1: The information room of Hospital A is only accessible to doctors, nurses, and hospital logistics staff, and no one else is allowed to enter. Among them, doctors and nurses are both wearing long white work uniforms, and logistics staff are both wearing short blue work tops and blue trousers. Therefore, the standard attributes of identities that meet the entry conditions of the information room of Hospital A can be preset as follows: long white work uniforms (standard attributes corresponding to the identities of doctors and nurses), short blue work tops and blue trousers (standard attributes corresponding to logistics staff). The standard attributes corresponding to the above three identities can be preset in the monitoring server of the information room of Hospital A, and the monitoring process can be as follows:
[0064] S1. Deploy several monitoring cameras at key positions in the information room of Hospital A that needs to be monitored, collect real-time images of the area to be monitored, and transmit the collected video images to the monitoring server of the information room of Hospital A.
[0065] S2. The monitoring server of the information room of Hospital A uses a pedestrian detection model to detect human targets that appear in the video image.
[0066] S3. The monitoring server of the information room of Hospital A uses a multi-attribute classification model to classify the relevant attributes of the human targets detected in the previous step, and obtains various attributes of the person. Among them, the various attributes of the person include whether wearing a hat, the color, texture, and style of the clothes, whether there is an epaulet, etc.
[0067] S4. Whitelist identity setting, adding doctors, nurses, and hospital logistics staff to the whitelist. Among them, doctors and nurses are defined as wearing long white work uniforms, and hospital logistics staff are defined as wearing short blue work tops and blue trousers. That is, add the standard attributes of identities that meet the entry conditions of the information room of Hospital A to the whitelist. In specific implementation, a blacklist for prohibiting entry into the information room of Hospital A can also be set according to actual needs, and this embodiment does not make specific limitations on this.
[0068] S5. Person identity matching: When the system discovers a target that does not match the identity in the whitelist, it will record an illegal intrusion event and issue an alarm, notifying relevant staff to verify the identity of the illegal intruder. That is, according to the various attributes of the person obtained in S3 and the standard attributes in the whitelist, identify whether the identity of the person entering the information room of Hospital A is a doctor, nurse, or hospital logistics staff of Hospital A.
[0069] Monitoring scenario 2: The reference room of Hospital B. Only doctors, nurses, and logistics staff are allowed to enter the reference room of Hospital B. Doctors may only wear long white work uniforms, while nurses will wear short white or pink work uniforms, and logistics staff wear short green tops and green trousers. Therefore, the standard attributes of identities that meet the entry conditions for the reference room of Hospital B can be preset as follows: long white work uniforms (standard attribute corresponding to doctors), short white or pink work uniforms (standard attribute corresponding to nurses), short green tops and green trousers (standard attribute corresponding to logistics staff). The standard attributes corresponding to the above three identities can be pre-stored in the monitoring server of the reference room of Hospital B, and the monitoring process can be as follows:
[0070] S1. Deploy a number of monitoring cameras at key positions in the reference room of Hospital B that needs to be monitored, collect real-time images of the area to be monitored, and transmit the collected video images to the monitoring server of the reference room of Hospital B.
[0071] S2. The monitoring server of the reference room of Hospital B uses a pedestrian detection model to detect the presence of human targets in the video image. After training the pedestrian detection model for Hospital A, this pedestrian detection model can be directly applied to Hospital B without retraining the pedestrian detection model.
[0072] S3. The monitoring server of the reference room of Hospital B uses a multi-attribute classification model to classify the relevant attributes of the human targets detected in the previous step to obtain multiple attributes of the person. Among them, the multiple attributes of the person include whether wearing a hat, the color, texture, and style of the clothes, and whether there are epaulets, etc. In specific implementation, when the multi-attribute classification model deployed in the reference room of Hospital A is trained, this multi-attribute classification model can be directly applied to the reference room of Hospital B without retraining the multi-attribute classification model.
[0073] S4. Whitelist identity setting, adding doctors, nurses, and hospital logistics staff to the whitelist. Among them, doctors are defined as long white work uniforms, nurses are defined as short white or pink work uniforms, and hospital logistics staff are defined as short green tops and green trousers. That is, add the standard attributes of identities that meet the entry conditions for the reference room of Hospital B to the whitelist. In specific implementation, a blacklist for prohibiting entry into the reference room of Hospital B can also be set according to actual needs, and this embodiment does not make specific limitations on this.
[0074] S5. Person identity matching: When the system discovers a target that does not match the identity in the whitelist, it will record an illegal intrusion event and issue an alarm, notifying relevant staff to verify the identity of the illegal intruder. That is, according to the multiple attributes of the person obtained in S3 and the standard attributes in the whitelist, identify whether the identity of the person entering the reference room of Hospital B is a doctor, nurse, or hospital logistics staff of Hospital B.
[0075] It should be noted that the above examples in this embodiment are all illustrative examples for easy understanding and do not limit the technical solutions of the present invention.
[0076] The beneficial effects of this embodiment are as follows: strong generalization performance, good flexibility, high efficiency, can effectively verify identities, improve the emergency response ability to illegal intrusion events, and are conducive to timely warning and prevention. It is mainly manifested in the following aspects:
[0077] 1. Compared with the models applicable to specific scenarios in the prior art, in this embodiment, identities that meet the entry conditions of the monitoring scenario are defined through standard attributes. Different monitoring scenarios can define different standard attributes, so that a multi-attribute classification model can be trained in this embodiment to adapt to the monitoring requirements of different monitoring scenarios. Therefore, the multi-attribute classification model in this embodiment does not need to retrain the network when migrating to other monitoring scenarios, has stronger generalization ability, can be flexibly applied to various monitoring scenarios, and is conducive to the large-scale deployment of the model.
[0078] 2. This embodiment can use a large number of publicly available image datasets to train the network model. In actual deployment scenarios, the workload of collecting data is large and the diversity of data is limited. This embodiment can use a large number of publicly available image datasets when training the multi-attribute classification model without having to collect data in actual deployment scenarios, which simplifies the cumbersome process of obtaining datasets and can use more data to train the multi-attribute classification model.
[0079] 3. The multi-attribute classification model used in this embodiment, that is, the multi-task classification network adopts the form of a shared backbone network, which can enable the network to learn more shared feature representations and improve the generalization effect of the network. Compared with Figure 2 training a model for each task as shown, this embodiment only uses one multi-attribute classification model, effectively improving the running efficiency of the network.
[0080] The second embodiment of this application relates to an identity recognition method. This embodiment is a further improvement of the first embodiment. The main improvement lies in: introducing an attention mechanism into the multi-attribute classification model, such as Figure 3As shown, after using a shared backbone network to extract features and obtain an intermediate feature map, when classifying the attributes of a certain area of a target person, a mask image corresponding to the area can be predicted first, and then the mask images corresponding to different areas are applied to the intermediate feature map to obtain target area feature maps corresponding to different areas in the intermediate feature map. Finally, various attributes of the target object are determined based on the target area feature maps corresponding to different areas. For example, when predicting the color of a person's upper garment, after using a shared backbone network to extract features and obtain an intermediate feature map, a mask image corresponding to the upper garment area can be predicted first, and then the mask image is applied to the intermediate feature map to remove the areas in the intermediate feature map that are irrelevant to the upper garment area. Finally, the color of the upper garment is predicted. The main improvements of this application are described below:
[0081] In this embodiment, it is equivalent to a further improvement on "determining various attributes of a target person according to a pre-trained multi-attribute classification model" in the first embodiment. The difference between the multi-attribute classification model in this embodiment and the multi-attribute classification model in the first embodiment lies in: the sample sets constructed when training the model are different. In the first embodiment, various attributes of the people in the images that meet the preset annotation conditions in the image dataset are annotated to construct a sample set; in this embodiment, various attributes of the people in the images that meet the preset annotation conditions in the image dataset and different areas of the people are annotated to construct a sample set. That is to say, in the first embodiment, various attributes of the people are annotated, and in this embodiment, in addition to annotating various attributes of the people, different areas of the people are also annotated.
[0082] In an example, the annotation of different areas of a person can refer to FIGS. 4(a) and 4(b). Among them, FIG. 4(a) is the original image without annotation, and in FIG. 4(b), the upper garment area, the pants area, and the hat area of the head are marked with different colors. In this embodiment, the implementation manner of "determining various attributes of a target person according to a pre-trained multi-attribute classification model" can be as Figure 5 shown, including:
[0083] Step 501: Input the video image into the backbone network in the multi-attribute classification model to obtain an intermediate feature map.
[0084] The backbone network in the multi-attribute classification model in this embodiment can be a Residual Neural Network (abbreviated as: ResNet), and ResNet can be further ResNet18. ResNet18 has fewer parameters and can achieve higher speed and accuracy. ResNet18 can extract the features of the video image to obtain the intermediate feature map corresponding to the video image.
[0085] Step 502: Determine the mask images corresponding to different regions of the target person in the intermediate feature map.
[0086] Specifically, after passing through several convolutional layers in the multi-attribute classification model, the intermediate feature map can obtain the mask images corresponding to different regions of the target person in the intermediate feature map. Among them, the mask image can be understood as a binary image. For example, the mask image corresponding to the upper garment region of the intermediate feature map can refer to Figure 6 , that is, the values within the upper garment region are all 1, and the values in the remaining regions are all 0.
[0087] Step 503: Apply the mask images corresponding to different regions to the intermediate feature map to obtain the target region feature maps corresponding to different regions in the intermediate feature map.
[0088] Step 504: Determine multiple attributes of the target object according to the target region feature maps corresponding to different regions respectively.
[0089] In an example, the intermediate feature map can be multiplied by the mask images corresponding to different regions respectively to obtain the target region feature maps corresponding to different regions in the intermediate feature map. Determine multiple attributes of the target object according to the target region feature maps corresponding to different regions respectively. By multiplying the intermediate feature map by the mask images corresponding to different regions respectively, information irrelevant to the currently concerned region can be removed, so that the attention of the network can be focused on the target region that needs to be focused on.
[0090] For example, when focusing on the relevant attributes of the upper garment region, the information in the image that does not belong to the upper garment region may affect the judgment of the network. Therefore, the intermediate feature map can be multiplied by the mask image corresponding to the upper garment region to remove the information irrelevant to the upper garment region, so that the attention of the network can be focused on the upper garment region that needs to be focused on, that is, obtain the target region feature map corresponding to the upper garment region. Then, according to the target region feature map corresponding to the upper garment region, determine the relevant attributes of the upper garment region of the target object. For example, according to the target region feature map corresponding to the upper garment region, determine the upper garment color and / or upper garment style of the target object.
[0091] For another example, when focusing on the relevant attributes of the pants region, the information in the image that does not belong to the pants region may affect the judgment of the network. Therefore, the intermediate feature map can be multiplied by the mask image corresponding to the pants region to remove the information irrelevant to the pants region, so that the attention of the network can be focused on the pants region that needs to be focused on, that is, obtain the target region feature map corresponding to the pants region. Then, according to the target region feature map corresponding to the pants region, determine the relevant attributes of the pants region of the target object. For example, according to the target region feature map corresponding to the pants region, determine the pants color and / or pants style of the target object.
[0092] In a specific implementation, determining various attributes of a target object based on the target region feature maps corresponding to different regions may include: determining the upper garment color and / or upper garment style of the target object according to the target region feature map corresponding to the upper garment region; determining the pants color and / or pants style of the target object according to the target region feature map corresponding to the pants region; determining whether the target object wears a hat and / or wears glasses according to the target region feature map corresponding to the head region, etc.
[0093] In this embodiment, by adding an attention mechanism, that is, when determining an attribute of a certain region of a target person, first determining a mask image of the region, applying the mask image of the region to the intermediate feature map to remove irrelevant background information, and then performing attribute classification of the region, the accuracy of determining various attributes of the target object can be effectively improved.
[0094] The step division of the above various methods is only for clear description. When implementing, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, they are all within the protection scope of this patent; making insignificant modifications to the algorithm or process or introducing insignificant designs, but not changing the core design of its algorithm and process are all within the protection scope of this patent.
[0095] The third embodiment of the present invention relates to a training method for a multi-attribute classification model, as Figure 7 shown, including:
[0096] Step 701: Obtain a publicly available image dataset.
[0097] Step 702: Label various attributes of the people in the images in the image dataset that meet the preset annotation conditions, and construct a sample set.
[0098] Step 703: Determine the structure of the network and configure the network hyperparameters of the network.
[0099] Step 704: Train the network configured with network hyperparameters according to the sample set to obtain a multi-attribute classification model.
[0100] It is not difficult to find that the implementation process of the training method for the multi-attribute classification model in this embodiment has been introduced in the first embodiment and the second embodiment. The relevant technical details mentioned in the first embodiment and the second embodiment are still valid in this embodiment. To avoid repetition, they are not described here again. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the first embodiment to the second embodiment.
[0101] In this embodiment, a large number of publicly available image datasets can be used for training the network model. In actual deployment scenarios, the workload of collecting data is large, and the diversity of the data is limited. In this embodiment, a large number of publicly available image datasets can be used when training the multi-attribute classification model, without having to collect data in the actual deployment scenario. This simplifies the cumbersome process of obtaining the dataset and can use more data to train the multi-attribute classification model. Moreover, the multi-attribute classification model used in this embodiment, that is, the multi-task classification network adopts the form of a shared backbone network, which can enable the network to learn more shared feature representations and improve the generalization effect of the network.
[0102] The fourth embodiment of the present invention relates to a training device for a multi-attribute classification model, as Figure 8 shown, including:
[0103] An acquisition module 801, configured to acquire a publicly available image dataset;
[0104] A labeling module 802, configured to label multiple attributes of the people in the images in the image dataset that meet the preset labeling conditions, and construct a sample set;
[0105] A configuration module 803, configured to determine the structure of the network and configure the network hyperparameters of the network;
[0106] A training module 804, configured to train the network configured with network hyperparameters according to the sample set to obtain a multi-attribute classification model.
[0107] It is not difficult to find that this embodiment is a device embodiment corresponding to the third embodiment. The relevant technical details and technical effects mentioned in the third embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied in the third embodiment.
[0108] The fifth embodiment of the present invention relates to an electronic device, as Figure 9 shown, including at least one processor 901; and, a memory 902 communicatively connected to the at least one processor 901; wherein, the memory 902 stores instructions executable by the at least one processor 901, and the instructions are executed by the at least one processor 901 to enable the at least one processor 901 to execute the identity recognition method in the first or second embodiment.
[0109] Among them, the memory 902 and the processor 901 are connected in a bus manner. The bus may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors 901 and the memory 902 together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, etc., which are well known in the art, so they will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single component or multiple components, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium. The data processed by the processor 901 is transmitted on the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor 901.
[0110] The processor 901 is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory 902 can be used to store the data used by the processor 901 when executing operations.
[0111] The sixth embodiment of the present invention relates to a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the above method embodiment is implemented.
[0112] That is, those skilled in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by instructing relevant hardware through a program. The program is stored in a storage medium, including several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0113] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present invention, and in practical applications, various changes can be made in form and details without departing from the spirit and scope of the present invention.
Claims
1. An identity recognition method, characterized in that, Including: Obtain video images within a monitoring scenario; If a target person appears in the detected video images, determine multiple attributes of the target person according to a pre-trained multi-attribute classification model; wherein, the multi-attribute classification model is trained according to a pre-constructed sample set, and the sample set includes a number of images labeled with attributes; Determine the standard attributes of identities that meet the entry conditions of the monitoring scenario, wherein the standard attributes include multiple standard attributes corresponding to multiple identities; Determine the priorities of the multiple standard attributes, wherein the priorities are determined based on the actual number of people corresponding to multiple identities in the monitoring scenario, and the higher the actual number of people corresponding to an identity, the higher the priority of the standard attributes corresponding to that identity; In accordance with the priorities of the multiple standard attributes, sequentially match the multiple attributes of the target person with each of the standard attributes; If the multiple attributes of the target person match the standard attributes corresponding to any one of the identities successfully, identify that the identity of the target person meets the entry conditions.
2. The identity recognition method according to claim 1, wherein The multi-attribute classification model is trained through the following training method: Obtain a publicly available image dataset; Label multiple attributes of the people in the images in the image dataset that meet the preset labeling conditions to construct the sample set; Determine the structure of the network and configure the network hyperparameters of the network; Train the network configured with the network hyperparameters according to the sample set to obtain the multi-attribute classification model.
3. The identity recognition method according to claim 2, wherein The labeling multiple attributes of the people in the images in the image dataset that meet the preset labeling conditions to construct the sample set includes: Label multiple attributes of the people in the images in the image dataset that meet the preset labeling conditions and different regions of the people to construct the sample set; The determining multiple attributes of the target person according to a pre-trained multi-attribute classification model includes: Input the video images into the backbone network in the multi-attribute classification model to obtain intermediate feature maps; Determine the mask images corresponding to different regions of the target person in the intermediate feature maps; Apply the mask images corresponding to the different regions to the intermediate feature maps to obtain target region feature maps corresponding to the different regions in the intermediate feature maps respectively; Determine multiple attributes of the target person according to the target region feature maps corresponding to the different regions respectively.
4. The identity recognition method according to claim 3, wherein The applying the mask images corresponding to the different regions to the intermediate feature maps to obtain target region feature maps corresponding to the different regions in the intermediate feature maps respectively includes: Multiply the intermediate feature maps with the mask images corresponding to the different regions respectively to obtain target region feature maps corresponding to the different regions in the intermediate feature maps respectively.
5. An electronic device, characterized in that, Including: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the identity recognition method according to any one of claims 1 to 4.
6. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the identity recognition method described in any one of claims 1 to 4.
Citation Information
Patent Citations
Image retrieval method and device
CN110134810A
Pedestrian attribute identification method based on background suppression
CN110222636A
Labor protection article wearing condition detection and identity recognition method based on deep learning
CN111488804A