Personnel detection method and device in business scene, and computer equipment
By using pre-trained personnel identification and classification models in business scenarios, combining personnel detection box information and regional characteristics, the problem of low personnel detection accuracy in the prior art is solved, and higher detection accuracy and automatic analysis capabilities are achieved.
Patent Information
- Application Number
- CN202510347976.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, personnel detection accuracy is not high, missed detection rates and missed detection rates are high, making it difficult to effectively deal with personnel behavior judgments in business scenarios.
By obtaining the original frame image of the target business scenario, inputting it to the pre-trained personnel identification model, obtaining personnel detection box information, and matching the pre-trained personnel classification model with the final personnel detection result.
It improves the accuracy of personnel detection in business scenarios, effectively improves the accuracy of detection of characters wearing different clothes, and helps the automatic analysis and positioning of personnel detection.
Smart Images

Figure CN120220191A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and particularly to a method, apparatus, computer device, computer-readable storage medium, and computer program product for detecting personnel in a business scenario. Background Art
[0002] With the development of computer vision technology, it is of great significance to judge the identity of a person through clothing for the discrimination and detection of the identity of a person in a business scenario.
[0003] Traditional detection methods usually extract useful features from preprocessed images, train a simple classification model using the extracted features, and then input the new images into the model after preprocessing and feature extraction steps to obtain the category output by the model as the prediction result. Since the detection accuracy of the categories of people wearing different clothes in image recognition is extremely important for the discrimination of personnel behavior in a business scenario, however, the false negative rate and false positive rate of the above methods are relatively high, and the processing effect is not good.
[0004] Therefore, there is a problem of low accuracy in personnel detection in the related art. Summary of the Invention
[0005] Based on this, it is necessary to provide a method, apparatus, computer device, computer-readable storage medium, and computer program product for detecting personnel in a business scenario that can improve the accuracy of personnel detection in view of the above technical problems.
[0006] In a first aspect, the present application provides a method for detecting personnel in a business scenario, the method comprising:
[0007] Obtain an original frame image collected for a target business scenario, and input the original frame image into a pre-trained personnel recognition model to obtain personnel detection box information; the pre-trained personnel recognition model is trained based on the pose data of different categories of personnel in different business scenarios;
[0008] According to the personnel detection box information, crop a personnel region image from the original frame image, perform feature extraction on the personnel region image to obtain personnel region features;
[0009] Input the personnel region features into a pre-trained personnel classification model to obtain personnel category prediction information; the pre-trained personnel classification model is trained based on the wearing characteristics of each category of personnel in the different business scenarios;
[0010] Combine the personnel detection box information and the personnel category prediction information to determine the personnel detection result of the original frame image; the personnel detection result is used to indicate the identity category and location corresponding to the clothing worn by different personnel in the target business scenario.
[0011] In one embodiment, inputting the original frame image into a pre-trained person recognition model to obtain person detection box information includes:
[0012] Processing the original frame image through the pre-trained person recognition model to obtain a model output result; the model output result includes bounding boxes, object labels, and recognition confidence levels corresponding to multiple recognition objects respectively; the object labels are used to distinguish whether the corresponding recognition object is a human body.
[0013] Performing detection box screening according to a pre-set confidence threshold, and the object labels and recognition confidence levels corresponding to each recognition object, and obtaining the person detection box information based on the bounding boxes of the recognition objects that are human bodies.
[0014] In one embodiment, the method further includes:
[0015] Collecting first person sample images containing various human postures based on the limb postures of different categories of personnel in multiple business scenarios.
[0016] Performing first data augmentation processing on the collected first person sample images to obtain person posture sample data; the first data augmentation processing includes data cleaning and data flipping; the person posture sample data includes a training sample set for model training and a validation sample set for model verification.
[0017] In one embodiment, the method further includes:
[0018] Training a first initial model according to the training sample set and a first loss function to obtain a model training result.
[0019] Performing loss calculation processing based on the validation sample set, determining an optimal model according to the model training result, and evaluating the model performance of the optimal model to obtain the pre-trained person recognition model.
[0020] In one embodiment, the method further includes:
[0021] For each category of personnel with different wearing characteristics in multiple business scenarios, obtaining second person sample images containing different wearing characteristics under different lights and different angles.
[0022] Performing second data augmentation processing on the obtained second person sample images to obtain person identity feature sample images; the second data augmentation processing includes geometric transformation, color space transformation, and noise addition.
[0023] In one embodiment, the method further includes:
[0024] Divide the sample image of the personnel identity characteristics into a pre-training data set and a fine-tuning data set;
[0025] Train the second initial model according to the pre-training data set and the second loss function to obtain a pre-training result;
[0026] Perform model fine-tuning processing by combining the pre-training result, the fine-tuning data set, and the preset training layer information to obtain the pre-trained personnel classification model.
[0027] In a second aspect, the present application also provides a personnel detection device in a business scenario. The device includes:
[0028] An image recognition module, configured to obtain an original frame image collected for a target business scenario, input the original frame image into a pre-trained personnel recognition model, and obtain personnel detection box information; the pre-trained personnel recognition model is trained based on the pose data of different categories of personnel in different business scenarios;
[0029] A personnel area feature extraction module, configured to crop a personnel area image from the original frame image according to the personnel detection box information, and perform feature extraction on the personnel area image to obtain personnel area features;
[0030] A personnel category prediction module, configured to input the personnel area features into a pre-trained personnel classification model to obtain personnel category prediction information; the pre-trained personnel classification model is trained based on the wearing characteristics of each category of personnel in the different business scenarios;
[0031] A personnel detection result obtaining module, configured to determine the personnel detection result of the original frame image by combining the personnel detection box information and the personnel category prediction information; the personnel detection result is used to indicate the identity category and location corresponding to the clothing of different personnel in the target business scenario.
[0032] In a third aspect, the present application also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.
[0033] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.
[0034] In a fifth aspect, the present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the above method are implemented.
[0035] In the above business scenario, the method, device, computer device, computer-readable storage medium, and computer program product for personnel detection obtain the original frame image collected for the target business scenario, input the original frame image into a pre-trained personnel recognition model to obtain personnel detection box information. The pre-trained personnel recognition model is trained based on the pose data of different categories of personnel in different business scenarios. According to the personnel detection box information, the personnel area image is cropped from the original frame image, and feature extraction is performed on the personnel area image to obtain personnel area features. Then, the personnel area features are input into a pre-trained personnel classification model to obtain personnel category prediction information. The pre-trained personnel classification model is trained based on the wearing characteristics of various categories of personnel in different business scenarios. Furthermore, by combining the personnel detection box information and the personnel category prediction information, the personnel detection result of the original frame image is determined. This personnel detection result is used to indicate the corresponding identity categories and positions of different personnel wearing in the target business scenario, achieving the optimization of personnel detection in the business scenario. Based on combining the personnel detection result in the first stage with the personnel classification result in the second stage to generate the final personnel detection result, it can effectively improve the detection accuracy for the detection task of the identities or categories of people wearing different clothes in the business scenario, and contribute to the automatic analysis and positioning of personnel detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0037] Figure 1 It is a schematic flowchart of a method for personnel detection in a business scenario in one embodiment;
[0038] Figure 2 It is a schematic diagram of the detection process for the identities or categories of people wearing different clothes in a business scenario in one embodiment;
[0039] Figure 3 It is a schematic flowchart of a method for personnel detection in a business scenario in another embodiment;
[0040] Figure 4 It is a structural block diagram of a device for personnel detection in a business scenario in one embodiment;
[0041] Figure 5 It is an internal structure diagram of a computer device in one embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] To make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0043] In an exemplary embodiment, as Figure 1 shown, a method for detecting personnel in a business scenario is provided. In this embodiment, an example is given where this method is applied to a terminal. It can be understood that this method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps 101 to step 104. Among them:
[0044] Step 101, obtain the original frame image collected for the target business scenario, and input the original frame image into a pre-trained personnel recognition model to obtain personnel detection box information.
[0045] Among them, the pre-trained personnel recognition model can be trained based on the pose data of different categories of personnel in different business scenarios.
[0046] As an example, the scenarios to which the method for detecting personnel in the business scenario of the present application can be applied include, but are not limited to: based on a security monitoring system, real-time monitoring of security personnel wearing specific uniforms and giving early warnings for their illegal operation behaviors; in a closed-book examination venue (such as invigilators wearing uniforms and examinees wearing civilian clothes), identifying the identity of personnel by detecting the clothes of personnel and judging and giving early warnings for illegal behaviors (such as examinees gathering can be judged as an examination violation, while invigilators and examinees gathering may be for rule clarification); in an evaluation scenario, the identity of personnel can be confirmed by the different uniform colors worn by personnel, so as to judge whether there are illegal behaviors in the close contact between people (such as the system stipulates that evaluation experts wearing civilian clothes are not allowed to whisper to each other, while staff wearing blue vests are allowed to discuss work systems with supervisors wearing red vests, etc.).
[0047] In practical applications, the current business scenario to be detected can be used as the target business scenario. As Figure 2 shown, the original frame image collected for the target business scenario can be obtained, and the original frame image is input into a pre-trained personnel recognition model (such as an object detection model). By using the model to perform human body detection on the original frame image, the bounding box of the human body, that is, the personnel detection box information, can be output.
[0048] In an alternative embodiment, by conducting on-site exploration of the business scenario in advance to observe and collect data on the working methods in the actual business scenario, it is possible to ensure that the sample atlas for model training covers people in various different angles and poses. For example, for the bid evaluation scenario, the human poses can include sitting, standing, lying prone, walking, leaning back, etc., and the bid evaluation room can also be divided into different specifications according to the actual business needs. Then, it is necessary to collect image samples of various human poses from different angles in bid evaluation rooms of different specifications; furthermore, a complete convolutional neural network based on Faster R-CNN can be selected as the model for object detection, and the collected image samples can be used for model training to generate a pre-trained object detection model.
[0049] Specifically, a dataset covering various human poses at different angles in the corresponding business scenario can be collected for annotation and used for the model training of object detection; then, the trained model (i.e., the pre-trained personnel recognition model) can be deployed to the information system and processed by obtaining frame images (i.e., original frame images) in real time.
[0050] Step 102: According to the personnel detection box information, crop the personnel region image from the original frame image, and perform feature extraction on the personnel region image to obtain the personnel region feature.
[0051] In a specific implementation, as Figure 2 shown, the human body region can be cropped from the frame image based on the human body bounding box, and then a convolutional neural network (CNN) can be used to perform feature extraction on the cropped human body region; exemplarily, by using the convolutional layer of the convolutional neural network to perform feature extraction on the cropped personnel region image, a method of multiple pooling can be adopted to extract more representative features, that is, the personnel region feature.
[0052] For example, a convolutional kernel can be used to slide over the human body region to calculate the dot product of the convolutional kernel and the local region to generate a feature map. Then, after the convolutional operation, a non-linear activation function (such as ReLU) can be applied to introduce non-linearity and enhance the expression ability of the model. Furthermore, low-level features (such as edges and textures) can be extracted based on the shallow convolutional layer, and more high-level semantic features (such as the shape and structure of objects) can be extracted based on the deep convolutional layer. By stacking multiple convolutional layers, more complex features of the human body region can be extracted.
[0053] In one example, by pre-detecting and examining the actual business scenarios, images of personnel in various working postures in each business scenario are collected, and data augmentation techniques can be used to increase the sample set. The model of the complete convolutional neural network based on Faster R-CNN is trained for object detection, enabling it to have the ability to recognize personnel in different business scenarios. Then, by deploying the model in an information system that obtains video streams from video devices such as cameras and intercepts them as frame images, the model can detect the frame images, output the bounding box of the human body, and then crop out the human body area and use the convolutional neural network to extract human body features.
[0054] Step 103: Input the personnel area features into a pre-trained personnel classification model to obtain personnel category prediction information.
[0055] Among them, the pre-trained personnel classification model can be trained based on the wearing characteristics of various categories of personnel in different business scenarios.
[0056] In practical applications, as Figure 2 shown, a pre-trained ResNet classification model (i.e., the pre-trained personnel classification model) can be used to classify the extracted personnel area features to determine whether the human body is wearing the target clothing, and the classification prediction result (i.e., the personnel category prediction information) can be output.
[0057] Optionally, in the actual business scenario, by classifying the identities of personnel represented by different wearing characteristics and collecting image sets of personnel with different wearing characteristics under various lights in the business scenario (including various angles and different human postures to avoid overfitting of the sample set), and then through data augmentation technology processing, the ResNet classification model can be pre-trained and fine-tuned to obtain the pre-trained personnel classification model, enabling it to have the classification ability for people wearing different clothing categories in this business scenario. Furthermore, the pre-trained ResNet classification model can be used to predict the category of the extracted features, and the category with the highest probability can be selected as the prediction result.
[0058] Step 104: Combine the personnel detection box information and the personnel category prediction information to determine the personnel detection result of the original frame image.
[0059] Among them, the personnel detection result can be used to indicate the corresponding identity categories and positions of different personnel wearing in the target business scenario.
[0060] After obtaining the information of the person detection box and the person category prediction information, class name mapping can be used to complete the visualization of the category prediction result based on the original frame image. Thus, based on the two-stage detection method, a visual model can be used to identify the human body, and a classification model can be used for feature classification. The detection results of the two stages can be combined to generate the final detection result (i.e., the person detection result), such as outputting the category and position of a person wearing a certain piece of clothing.
[0061] Compared with the traditional method, the technical solution of this embodiment adopts a two-stage detection method for the identity of a person wearing specific clothing in a business scenario. By using the dataset of various human postures in the business scenario to train the visual model and making full use of the complete convolutional neural network visual model, human body detection is performed on the business scenario image to generate the bounding box of the human body. Then, by extracting features from the human body region and using the pre-trained classification model to classify the extracted features, it is determined whether the human body is wearing the target clothing. Furthermore, the classification result of the second stage is combined with the human body detection result of the first stage to output the final detection result, such as the position and category of a person wearing a certain piece of clothing, to complete the detection. Therefore, it can effectively improve the detection accuracy of the identity or category of people wearing different clothes in the business scenario, providing a theoretical basis and technical guidance for the accurate detection of people with different identities represented by different clothes in the business scenario in actual projects.
[0062] In the above-mentioned person detection method in the business scenario, by acquiring the original frame image collected for the target business scenario, inputting the original frame image into the pre-trained person recognition model to obtain the person detection box information, cropping the person region image from the original frame image according to the person detection box information, extracting features from the person region image to obtain the person region features, and then inputting the person region features into the pre-trained person classification model to obtain the person category prediction information. Furthermore, by combining the person detection box information and the person category prediction information, the person detection result of the original frame image is determined, realizing the optimization of person detection in the business scenario. Based on combining the person detection result of the first stage with the person classification result of the second stage to generate the final person detection result, it can effectively improve the detection accuracy for the detection task of the identity or category of people wearing different clothes in the business scenario, and contribute to the automatic analysis and positioning of person detection.
[0063] In an exemplary embodiment, the step of inputting the original frame image into the pre-trained person recognition model to obtain the person detection box information may include the following steps:
[0064] Process the original frame image through the pre-trained person recognition model to obtain a model output result; the model output result includes bounding boxes, object labels, and recognition confidence levels corresponding to each of the multiple recognition objects; the object labels are used to distinguish whether the corresponding recognition object is a human body; according to a preset confidence threshold, and the object labels and recognition confidence levels corresponding to each recognition object, perform detection box screening, and obtain the person detection box information based on the bounding boxes of the recognition objects that are human bodies.
[0065] In practical applications, a pre-trained Faster R-CNN model (i.e., the pre-trained person recognition model) can be loaded, the model can be set to the evaluation mode and moved to the information system device, and then by performing normalization processing on the original frame image, the pre-processed image can be input into the model, and the model can output detection results including bounding boxes, class labels, and confidence levels (i.e., the model output result). Furthermore, the detection boxes can be screened according to the class labels (e.g., 0 represents "other" and 1 represents "person") and the confidence threshold, and the detection boxes can be drawn on the image, and the detection results can be displayed in the form of an [x1, y1, x2, y2] array, so as to further crop the human body region from the original frame image according to the coordinates of the bounding boxes.
[0066] In this embodiment, by processing the original frame image through the pre-trained person recognition model to obtain a model output result, and then performing detection box screening according to the preset confidence threshold, and the object labels and recognition confidence levels corresponding to each recognition object, and obtaining the person detection box information based on the bounding boxes of the recognition objects that are human bodies, the person detection box can be accurately extracted, effectively improving the accuracy and efficiency of person detection.
[0067] In an exemplary embodiment, the following steps may further be included:
[0068] Based on the body postures of different types of personnel in multiple business scenarios, collect first person sample images containing various human postures; through performing first data augmentation processing on the collected first person sample images, obtain person posture sample data; the first data augmentation processing includes data cleaning and data flipping; the person posture sample data includes a training sample set for model training and a validation sample set for model validation.
[0069] In one example, for the real business scenario of distinguishing personnel identity categories through personnel clothing, to better adapt to the processing scenario, considering that various types of personnel will present different body postures. For example, in the scenarios of examinations and bid evaluations, teachers, examinees, bid evaluation personnel, and staff may all have body postures such as sitting, lying back, standing, walking, and talking. In the work safety supervision scenario, staff may also have body postures such as lying prone (repairing equipment), etc. Therefore, when collecting data, the possible body postures of personnel can be analyzed according to the characteristics of each business scenario, so as to collect human image samples in various postures (i.e., the first personnel sample images), and through data enhancement methods such as data cleaning and flipping (i.e., the first data enhancement processing), the number of the sample set can be increased, and the data can be divided into two categories: the training set and the validation set, namely the training sample set and the validation sample set.
[0070] In this embodiment, by collecting the first personnel sample images containing various human body postures based on the body postures of different categories of personnel in multiple business scenarios, and then performing the first data enhancement processing on the collected first personnel sample images to obtain the personnel posture sample data, the diversity and quality of the sample data can be effectively improved, which helps to enhance the recognition and generalization ability of the model in complex environments.
[0071] In an exemplary embodiment, the following steps may further be included:
[0072] According to the training sample set and the first loss function, model training is performed on the first initial model to obtain the model training result; through loss calculation processing based on the validation sample set, the optimal model is determined according to the model training result, and the model performance of the optimal model is evaluated to obtain the pre-trained personnel recognition model.
[0073] In specific implementation, a pre-trained convolutional network (such as ResNet, VGG) can be used as the backbone network (i.e., the first initial model) to extract image features; by encapsulating the labeled images and bounding box information into a dataset class and performing batch loading, the backbone network can be used to extract features, the RPN (Region Proposal Network) can be used to generate candidate regions, and the features of the candidate regions can be extracted through the ROI pooling layer. Based on the defined loss function (i.e., the first loss function), the loss in training is calculated, and the learning rate can be adjusted according to the training progress to improve the training efficiency and accelerate convergence, so as to avoid falling into the local optimal solution.
[0074] Optionally, during the training process, the model weights can be saved regularly, the optimal model can be selected according to the loss of the data validation set, and the model performance can be evaluated by calculating metrics such as mean average precision (mAP) and recall rate. Furthermore, a pre-trained object detection model based on Faster R-CNN, that is, a pre-trained personnel recognition model, can be obtained.
[0075] In this embodiment, by training the first initial model according to the training sample set and the first loss function, the model training result is obtained. Then, through loss calculation processing based on the validation sample set, the optimal model is determined according to the model training result, and the model performance of the optimal model is evaluated to obtain a pre-trained personnel recognition model, which can help improve the model performance and recognition accuracy.
[0076] In an exemplary embodiment, the following steps may further be included:
[0077] For each category of personnel with different wearing characteristics in the multiple business scenarios, obtain second personnel sample images containing different wearing characteristics under different lights and different angles; through second data augmentation processing on the obtained second personnel sample images, obtain personnel identity feature sample images; the second data augmentation processing includes geometric transformation, color space transformation, and noise addition.
[0078] In one example, by collecting images of personnel with different wearing characteristics in the real business scenario, such as considering different angles under cloudy days, sunny days, daytime, evening, lighting, and natural light during collection, as well as image sampling of the human posture, the second personnel sample images are obtained, thereby ensuring that various samples are sufficient and the feature differences are obvious; data augmentation methods such as data geometric transformation, color space transformation, and noise addition (i.e., the second data augmentation processing) can be used to maintain the features of the original categories and increase the number of the sample set, and overfitting images can be eliminated through inspection, and finally, personnel identity feature sample images are obtained.
[0079] In this embodiment, by obtaining second personnel sample images containing different wearing characteristics under different lights and different angles for each category of personnel with different wearing characteristics in multiple business scenarios, and then through second data augmentation processing on the obtained second personnel sample images to obtain personnel identity feature sample images, the diversity and robustness of the samples can be effectively improved, and the adaptability and accuracy of identity recognition in different environments are enhanced.
[0080] In an exemplary embodiment, the following steps may further be included:
[0081] Divide the sample image of the personnel identity characteristics into a pre-training data set and a fine-tuning data set; perform model training on the second initial model according to the pre-training data set and the second loss function to obtain a pre-training result; combine the pre-training result, the fine-tuning data set, and the preset training layer information to perform model fine-tuning processing to obtain the pre-trained personnel classification model.
[0082] Specifically, the size of the enhanced image (i.e., the sample image of the personnel identity characteristics) can be set to 224×224, and the image can be normalized using the mean and standard deviation of the data set. The image data format can be converted, and the image data can be divided into a pre-training data set and a fine-tuning data set through image features. For example, the data sets can be distinguished according to the proportion of the wearing area. If the clothing accounts for more than 1 / 5 of the image and has obvious clothing features such as logos and textures, it can be classified into the fine-tuning data set, and the others can be classified into the pre-training data set.
[0083] By loading and batch-processing the pre-training data, and loading the pre-trained ResNet model (i.e., the second initial model), the fully connected layer of ResNet can be modified according to the number of categories in the data set (for example, if there are three types of personnel wearing identity images, the number of layers can be set to 3), and the optimizer and loss function (i.e., the second loss function) can be defined for model pre-training to obtain the pre-training result.
[0084] Furthermore, some layers (such as the early layers) can be selected to be frozen, and only the last three layers can be trained (if the data set is large, the number of training layers can be appropriately increased), that is, the preset training layer information. By training the model using the fine-tuning data set, it can be made to have the classification ability for people wearing different clothes in the business scenario, and the pre-trained personnel classification model can be obtained.
[0085] In this embodiment, by dividing the sample image of the personnel identity characteristics into a pre-training data set and a fine-tuning data set, then performing model training on the second initial model according to the pre-training data set and the second loss function to obtain a pre-training result, and further combining the pre-training result, the fine-tuning data set, and the preset training layer information to perform model fine-tuning processing to obtain the pre-trained personnel classification model, the model can be trained and fine-tuned efficiently, and the classification accuracy and generalization ability of the model are improved.
[0086] In an exemplary embodiment, as Figure 3 shown, a schematic flow diagram of another method for personnel detection in a business scenario is provided. In this embodiment, the method includes the following steps:
[0087] In step 301, based on the body postures of different categories of personnel in multiple business scenarios, the first personnel sample images containing various human postures are collected for the first data augmentation process to obtain personnel posture sample data. In step 302, according to the training sample set and the first loss function, the first initial model is trained to obtain the model training result. By performing loss calculation processing based on the validation sample set, the optimal model is determined according to the model training result, and the model performance of the optimal model is evaluated to obtain the pre-trained personnel recognition model. In step 303, for different categories of personnel with different wearing characteristics in multiple business scenarios, the second personnel sample images containing different wearing characteristics under different lights and different angles are obtained for the second data augmentation process to obtain personnel identity feature sample images. In step 304, the personnel identity feature sample images are divided into a pre-training data set and a fine-tuning data set. According to the pre-training data set and the second loss function, the second initial model is trained to obtain the pre-training result. In step 305, combined with the pre-training result, the fine-tuning data set, and the preset training layer information, model fine-tuning processing is performed to obtain the pre-trained personnel classification model. In step 306, the original frame image collected for the target business scenario is obtained, and the original frame image is input into the pre-trained personnel recognition model to obtain the personnel detection box information. The personnel region image is cropped from the original frame image, and feature extraction is performed on the personnel region image to obtain the personnel region features. In step 307, the personnel region features are input into the pre-trained personnel classification model to obtain the personnel category prediction information. Combining the personnel detection box information and the personnel category prediction information, the personnel detection result of the original frame image is determined.
[0088] It should be noted that the specific limitations of the above steps can refer to the specific limitations of a personnel detection method in a business scenario described above, and will not be elaborated here.
[0089] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.
[0090] Based on the same inventive concept, an embodiment of the present application further provides a personnel detection device in a service scenario for implementing the personnel detection method in the above-mentioned service scenario. The implementation solutions provided by this device to solve problems are similar to those described in the above method. Therefore, the specific limitations in one or more embodiments of the personnel detection device in the service scenario provided below can refer to the limitations on the personnel detection method in the above text, and will not be elaborated here.
[0091] In an exemplary embodiment, as Figure 4 shown, a personnel detection device in a service scenario is provided, including:
[0092] An image recognition module 401, configured to obtain an original frame image collected for a target service scenario, input the original frame image into a pre-trained personnel recognition model, and obtain personnel detection box information; the pre-trained personnel recognition model is trained based on the pose data of different categories of personnel in different service scenarios;
[0093] A personnel area feature extraction module 402, configured to crop a personnel area image from the original frame image according to the personnel detection box information, perform feature extraction on the personnel area image, and obtain personnel area features;
[0094] A personnel category prediction module 403, configured to input the personnel area features into a pre-trained personnel classification model, and obtain personnel category prediction information; the pre-trained personnel classification model is trained based on the wearing characteristics of each category of personnel in the different service scenarios;
[0095] A personnel detection result obtaining module 404, configured to combine the personnel detection box information and the personnel category prediction information to determine the personnel detection result of the original frame image; the personnel detection result is used to indicate the identity categories and locations corresponding to the different personnel's wearings in the target service scenario.
[0096] In one embodiment, the image recognition module 401 is specifically configured to process the original frame image through the pre-trained personnel recognition model to obtain a model output result; the model output result includes bounding boxes, object labels, and recognition confidences respectively corresponding to multiple recognition objects; the object labels are used to distinguish whether the corresponding recognition object is a human body; according to a preset confidence threshold, and the object labels and recognition confidences corresponding to each recognition object, perform detection box screening, and obtain the personnel detection box information based on the bounding boxes of the recognition objects that are human bodies.
[0097] In one embodiment, the device further includes:
[0098] The first sample image acquisition module is used to collect the first personnel sample images containing various human postures based on the body postures of different types of personnel in multiple business scenarios;
[0099] The posture sample data acquisition module is used to obtain personnel posture sample data by performing first data augmentation processing on the collected first personnel sample images; the first data augmentation processing includes data cleaning and data flipping; the personnel posture sample data includes a training sample set for model training and a validation sample set for model validation.
[0100] In one embodiment, the device further includes:
[0101] The first initial model training module is used to perform model training on the first initial model according to the training sample set and the first loss function to obtain a model training result;
[0102] The personnel recognition model acquisition module is used to determine the optimal model based on the model training result by performing loss calculation processing based on the validation sample set, and evaluate the model performance of the optimal model to obtain the pre-trained personnel recognition model.
[0103] In one embodiment, the device further includes:
[0104] The second sample image acquisition module is used to obtain second personnel sample images containing different dressing features under different lights and different angles for different types of personnel with different dressing features in multiple business scenarios;
[0105] The feature sample image acquisition module is used to obtain personnel identity feature sample images by performing second data augmentation processing on the obtained second personnel sample images; the second data augmentation processing includes geometric transformation, color space transformation, and noise addition.
[0106] In one embodiment, the device further includes:
[0107] The sample image division module is used to divide the personnel identity feature sample images into a pre-training data set and a fine-tuning data set;
[0108] The second initial model training module is used to perform model training on the second initial model according to the pre-training data set and the second loss function to obtain a pre-training result;
[0109] The personnel classification model acquisition module is used to perform model fine-tuning processing by combining the pre-training result, the fine-tuning data set, and the preset training layer information to obtain the pre-trained personnel classification model.
[0110] In the above business scenario, each module in the personnel detection device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0111] In an exemplary embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 5 shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used for exchanging information between the processor and external devices. The communication interface of the computer device is used for communicating with external terminals in a wired or wireless manner. The wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a personnel detection method in a business scenario. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0112] Those skilled in the art can understand that Figure 5 the structure shown in
[0113] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0114] Obtain the original frame image collected for the target business scenario, input the original frame image into a pre-trained person recognition model, and obtain person detection box information; the pre-trained person recognition model is trained based on the pose data of different categories of people in different business scenarios;
[0115] According to the person detection box information, crop the person area image from the original frame image, perform feature extraction on the person area image, and obtain person area features;
[0116] Input the person area features into a pre-trained person classification model to obtain person category prediction information; the pre-trained person classification model is trained based on the wearing characteristics of each category of people in the different business scenarios;
[0117] Combine the person detection box information and the person category prediction information to determine the person detection result of the original frame image; the person detection result is used to indicate the identity categories and locations corresponding to the clothing of different people in the target business scenario.
[0118] In one embodiment, when the processor executes the computer program, it also implements the steps of the method for detecting people in the business scenario in the above-mentioned other embodiments.
[0119] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0120] Obtain the original frame image collected for the target business scenario, input the original frame image into a pre-trained person recognition model, and obtain person detection box information; the pre-trained person recognition model is trained based on the pose data of different categories of people in different business scenarios;
[0121] According to the person detection box information, crop the person area image from the original frame image, perform feature extraction on the person area image, and obtain person area features;
[0122] Input the person area features into a pre-trained person classification model to obtain person category prediction information; the pre-trained person classification model is trained based on the wearing characteristics of each category of people in the different business scenarios;
[0123] Combine the person detection box information and the person category prediction information to determine the person detection result of the original frame image; the person detection result is used to indicate the identity categories and locations corresponding to the clothing of different people in the target business scenario.
[0124] In one embodiment, when the computer program is executed by a processor, it also implements the steps of the personnel detection method in the business scenarios of the above other embodiments.
[0125] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the following steps:
[0126] Obtain the original frame image collected for the target business scenario, input the original frame image into a pre-trained personnel recognition model to obtain personnel detection box information; the pre-trained personnel recognition model is trained based on the pose data of different categories of personnel in different business scenarios;
[0127] According to the personnel detection box information, crop the personnel area image from the original frame image, perform feature extraction on the personnel area image to obtain personnel area features;
[0128] Input the personnel area features into a pre-trained personnel classification model to obtain personnel category prediction information; the pre-trained personnel classification model is trained based on the wearing characteristics of each category of personnel in the different business scenarios;
[0129] Combine the personnel detection box information and the personnel category prediction information to determine the personnel detection result of the original frame image; the personnel detection result is used to indicate the identity categories and positions of different personnel wearing in the target business scenario.
[0130] In one embodiment, when the computer program is executed by a processor, it also implements the steps of the personnel detection method in the business scenarios of the above other embodiments.
[0131] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0132] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0133] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded in the present application.
[0134] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A method for detecting people in a business scenario, characterized in that: The method comprises: Acquire an original frame image collected for a target business scene, input the original frame image into a pre-trained personnel recognition model, and obtain personnel detection frame information; the pre-trained personnel recognition model is trained based on posture data of different categories of personnel in different business scenes; According to the personnel detection frame information, a personnel region image is cropped from the original frame image, and features are extracted from the personnel region image to obtain personnel region features; Inputting the personnel area features into a pre-trained personnel classification model to obtain personnel category prediction information; the pre-trained personnel classification model is trained based on the clothing features of the personnel of each category in the different business scenarios; The personnel detection result of the original frame image is determined by combining the personnel detection frame information and the personnel category prediction information; the personnel detection result is used to indicate the identity category and location corresponding to the clothing worn by different personnel in the target business scene.
2. The method according to claim 1, characterized in that The inputting the original frame image into a pre-trained person recognition model to obtain person detection frame information includes: Processing the original frame image through the pre-trained person recognition model to obtain a model output result; the model output result includes a bounding box, an object label, and a recognition confidence corresponding to each of the multiple recognition objects; the object label is used to distinguish whether the corresponding recognition object is a human body; The detection frame is screened according to a preset confidence threshold, an object label corresponding to each of the identified objects, and a recognition confidence, and the person detection frame information is obtained based on the boundary frame of the identified object being a human body.
3. The method according to claim 1, characterized in that The method further comprises: Based on the body shapes of the different categories of people in the multiple business scenarios, collecting first person sample images containing various human body postures; Personnel posture sample data is obtained by performing a first data enhancement process on the collected first person sample image; the first data enhancement process includes data cleaning and data flipping; the person posture sample data includes a training sample set for model training and a verification sample set for model verification.
4. The method according to claim 3, characterized in that The method further comprises: Performing model training on the first initial model according to the training sample set and the first loss function to obtain a model training result; The pre-trained personnel recognition model is obtained by performing loss calculation processing based on the verification sample set, determining the optimal model according to the model training result, and evaluating the model performance of the optimal model.
5. The method according to claim 1, characterized in that The method further comprises: For each of the categories of persons with different clothing characteristics in the multiple business scenarios, obtaining second person sample images containing different clothing characteristics under different light and different angles; By performing a second data enhancement process on the acquired second person sample image, a person identity feature sample image is obtained; the second data enhancement process includes geometric transformation, color space transformation, and noise addition.
6. The method according to claim 5, characterized in that The method further comprises: Dividing the sample images of the person identity features into a pre-training data set and a fine-tuning data set; Performing model training on the second initial model according to the pre-training data set and the second loss function to obtain a pre-training result; The pre-training result, the fine-tuning data set, and the preset training layer information are combined to perform model fine-tuning processing to obtain the pre-trained personnel classification model.
7. A person detection device in a business scenario, characterized in that: The device comprises: An image recognition module is used to obtain an original frame image collected for a target business scene, input the original frame image into a pre-trained personnel recognition model, and obtain personnel detection frame information; the pre-trained personnel recognition model is trained based on posture data of different categories of personnel in different business scenes; A personnel region feature extraction module, used to crop a personnel region image from the original frame image according to the personnel detection frame information, perform feature extraction on the personnel region image, and obtain personnel region features; A personnel category prediction module, used for inputting the personnel area features into a pre-trained personnel classification model to obtain personnel category prediction information; the pre-trained personnel classification model is trained based on the clothing features of the personnel of each category in the different business scenarios; The personnel detection result obtaining module is used to determine the personnel detection result of the original frame image by combining the personnel detection frame information and the personnel category prediction information; the personnel detection result is used to indicate the identity category and location corresponding to the clothing worn by different personnel in the target business scene.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.