A detection method, device, equipment and medium for wearable target objects
A multi-model detection system with boundary jittering enhances the accuracy of identifying worn items by addressing false positives in complex scenes, improving the precision of detecting headwear or gloves.
Patent Information
- Application Number
- CN202210347919.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-01
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-04-01
AI Technical Summary
Existing wearable object detection models are prone to false alarms in complex scenarios, resulting in inaccurate detection results.
The multi-task network method is used to verify the image multiple times through three different models, including the first model judging whether the target part is worn by the target object, the second model judging the attribute category of the target object, and the third model judging whether the image is a difficult image, combining multiple boundary jitters and image verification to improve detection accuracy.
It significantly improves the accuracy of wearable target detection, reduces false alarms, and ensures the reliability of detection results.
Smart Images

Figure CN114782849B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of security monitoring technology, and particularly relates to a method, device, equipment and medium for detecting wearable objects. Background Art
[0002] With the popularization of intelligent technology, all walks of life have strict requirements for the wearing behaviors of staff. For example, kitchen staff are required to wear chef hats, sanitary gloves, etc., and workshop operators are required to wear workshop hats, safety gloves, etc. By standardizing the wearing behaviors of staff, the kitchen hygiene quality, workshop production safety quality, etc. can be effectively improved, thereby improving the service quality of various industries and enhancing the core competitiveness of enterprises. Therefore, it is crucial to accurately detect whether personnel in a specific place wear target objects such as headgear and gloves.
[0003] Currently, mainly a trained single model is used to detect images collected in a specific place to determine whether the personnel in the image wear target objects. However, in the face of interference from complex scenes, a single detection model usually has false alarms, and the obtained detection results of wearable objects are inaccurate. Summary of the Invention
[0004] Embodiments of this application provide a method, device, equipment and medium for detecting wearable objects, which are used to improve the accuracy of the detection results of wearable objects.
[0005] In a first aspect, this application provides a method for detecting wearable objects, including:
[0006] Obtain a first image including a target part of an object to be detected;
[0007] Input the first image into a first model to obtain a first result, where the first result represents whether the target part wears a target object; the first model is obtained by training with a first sample set, and the first sample set includes images marked for whether the target parts of historical objects wear target objects;
[0008] In response to the first result indicating that the target part does not wear a target object, input the first image into a second model and a third model to obtain a second result and a third result; where the second result represents the attribute category of the target object in the first image, the third result represents whether the first image is a difficult example image, and the attribute category of the target object in the difficult example image is a first attribute category; the second model is obtained by training with a second sample set, and the second sample set includes images marked for target objects of different attribute categories; the third model is obtained by training with a third sample set, and the third sample set includes images marked for whether the sample images are difficult example images;
[0009] Based on the second result and the third result, determine whether the target object is worn on the target part in the first image.
[0010] In a possible embodiment, based on the second result and the third result, determining whether the target object is worn on the target part in the first image includes:
[0011] In response to the second result characterizing that the attribute category of the target object in the first image is the second attribute category, and the third result characterizing that the first image is not a difficult example image, determine that the target object is not worn on the target part in the first image.
[0012] In a possible embodiment, the different attribute categories include a first attribute category in which the color of the target object is the target color, a second attribute category in which the target object is not worn on the target part, and a third attribute category in which the color of the target object is a non-target color.
[0013] In a possible embodiment, the target object is a headdress, and the target color is any one of black, white, or gray.
[0014] In a possible embodiment, obtaining a first image including a target part of a to-be-detected object includes:
[0015] Collect a target image of the to-be-detected object in a target environment;
[0016] If it is determined that there is a target part in the target image, based on a preset first rectangular frame, extract an image of the target part from the target image as the first image.
[0017] In a possible embodiment, in response to the first result characterizing that the target object is not worn on the target part, inputting the first image into a second model and a third model to obtain a second result and a third result includes:
[0018] Perform N times of boundary jittering on the first rectangular frame, and according to the position information of the first rectangular frame in the target image, obtain the position information of N second rectangular frames in the target image; where N is an odd number greater than or equal to 3, and each boundary jittering means expanding or contracting the four boundaries of the first rectangular frame respectively;
[0019] According to the position information of the N second rectangular frames in the target image, extract N second images from the target image;
[0020] Input each second image into the second model and the third model to obtain a second result and a third result corresponding to each second image.
[0021] In a possible embodiment, determining whether the target part in the first image wears a target object based on the second result and the third result includes:
[0022] If the corresponding second result indicates that the attribute category of the target object in the first image is the second attribute category, and the number of corresponding third results indicating that the first image is not a difficult example image is greater than or equal to a preset value, it is determined that the target part in the first image does not wear the target object; wherein, the preset value is proportional to the value of N.
[0023] In a possible embodiment, the first model is the first branch of a multi-task network, the second model is the second branch of the multi-task network, the third model is the third branch of the multi-task network, and the multi-task network further includes a backbone module;
[0024] Inputting the first image into the first model to obtain a first result includes:
[0025] Inputting the first image into the backbone module to obtain the image features of the first image;
[0026] Inputting the image features into the first branch to obtain a first result;
[0027] In response to the first result indicating that the target part does not wear the target object, inputting the first image into the second model and the third model to obtain a second result and a third result includes:
[0028] In response to the first result indicating that the target part does not wear the target object, inputting the image features into the second branch to obtain a second result, and inputting the image features into the third branch to obtain a third result.
[0029] In a second aspect, the present application provides a detection device for wearing a target object, including:
[0030] An acquisition module, configured to acquire a first image including the target part of the object to be detected;
[0031] An obtaining module, configured to input the first image into a first model to obtain a first result, where the first result indicates whether the target part wears a target object; the first model is obtained by training with a first sample set, and the first sample set includes images marked for whether the target part of a historical object wears a target object;
[0032] The obtaining module is further configured to, in response to the first result indicating that the target object is not worn on the target part, input the first image into a second model and a third model to obtain a second result and a third result; wherein, the second result indicates the attribute category of the target object in the first image, the third result indicates whether the first image is a difficult example image, and the attribute category of the target object in the difficult example image is a first attribute category; the second model is obtained by training with a second sample set, and the second sample set includes images labeled for target objects of different attribute categories; the third model is obtained by training with a third sample set, and the third sample set includes images labeled for whether the sample images are difficult example images;
[0033] A determining module, configured to determine whether the target object is worn on the target part in the first image based on the second result and the third result.
[0034] In a possible embodiment, the determining module is specifically configured to:
[0035] In response to the second result indicating that the attribute category of the target object in the first image is a second attribute category, and the third result indicating that the first image is not a difficult example image, determine that the target object is not worn on the target part in the first image.
[0036] In a possible embodiment, the different attribute categories include a first attribute category in which the color of the target object is the target color, a second attribute category in which the target object is not worn on the target part, and a third attribute category in which the color of the target object is a non-target color.
[0037] In a possible embodiment, the target object is a headdress, and the target color is any one of black, white, or gray.
[0038] In a possible embodiment, the obtaining module is specifically configured to:
[0039] Collect a target image of the object to be detected in the target environment;
[0040] If it is determined that there is a target part in the target image, then based on a preset first rectangular frame, extract the image of the target part from the target image as the first image.
[0041] In a possible embodiment, the obtaining module is specifically configured to:
[0042] Perform N times of boundary jitter on the first rectangular frame, and obtain the position information of N second rectangular frames in the target image according to the position information of the first rectangular frame in the target image; where N is an odd number greater than or equal to 3, and each boundary jitter means expanding or contracting the four boundaries of the first rectangular frame respectively;
[0043] Extract N second images from the target image according to the position information of the N second rectangular frames in the target image;
[0044] Input each second image into the second model and the third model to obtain a second result and a third result corresponding to each second image.
[0045] In a possible embodiment, the determining module is specifically configured to:
[0046] If the corresponding second result indicates that the attribute category of the target object in the first image is the second attribute category, and the number of second images corresponding to the third result indicating that the first image is not a difficult example image is greater than or equal to a preset value, it is determined that the target part in the first image is not wearing the target object; wherein, the preset value is proportional to the value of N.
[0047] In a possible embodiment, the first model is the first branch of a multi-task network, the second model is the second branch of the multi-task network, the third model is the third branch of the multi-task network, and the multi-task network further includes a backbone module;
[0048] The obtaining module is specifically configured to:
[0049] Input the first image into the backbone module to obtain the image features of the first image;
[0050] Input the image features into the first branch to obtain a first result;
[0051] In response to the first result indicating that the target part is not wearing the target object, input the image features into the second branch to obtain a second result, and input the image features into the third branch to obtain a third result.
[0052] In a third aspect, the present application provides a detection device for a wearable target object, including:
[0053] A memory for storing program instructions;
[0054] A processor for calling the program instructions stored in the memory and executing the method according to any one of the first aspects according to the obtained program instructions.
[0055] In a fourth aspect, the present application provides a computer-readable storage medium, the computer-readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by a computer, the computer is caused to execute the method according to any one of the first aspects.
[0056] In an embodiment of the present application, a first image including a target part is input into a first model. The first model is obtained by training with images marked for whether a target object is worn on the target part of a historical object. If the first result output by the first model indicates that the target object is not worn on the target part, the first image is input into a second model and a third model respectively to further verify the first result. Based on the second result output by the second model and the third result output by the third model, it is determined whether the target object is worn on the target part in the first image. In the embodiment of the present application, the first model, the second model, and the third model are trained with different sample sets. By combining the results output by the three different models, it is possible to more accurately determine whether the target part in the first image wears a headdress, improving the accuracy of detecting whether the target object is worn. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on the provided drawings without creative efforts.
[0058] Figure 1 FIG. is a schematic diagram of an application scenario of a method for detecting whether a target object is worn provided by an embodiment of the present application;
[0059] Figure 2 FIG. is a flow chart of a method for detecting whether a target object is worn provided by an embodiment of the present application Figure 1 ;
[0060] Figure 3 FIG. is a structural diagram of a multi-task network provided by an embodiment of the present application;
[0061] Figure 4 FIG. is a flow chart of a method for detecting whether a target object is worn provided by an embodiment of the present application Figure 2 ;
[0062] Figure 5 FIG. is a structural diagram of a device for detecting whether a target object is worn provided by an embodiment of the present application;
[0063] Figure 6 FIG. is a structural diagram of a device for detecting whether a target object is worn provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] To make the objectives, technical solutions and advantages of the present invention more clearly understood, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application. Without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other arbitrarily. And although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.
[0065] In the description and claims of the present application and the above accompanying drawings, the terms "first" and "second" are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0066] In the embodiments of the present application, "a plurality" may represent at least two, for example, it may be two, three or more, and the embodiments of the present application do not make limitations.
[0067] To improve the accuracy of the detection result of the wearable target, an embodiment of the present application provides a detection method for the wearable target, and this method can be executed by a detection device for the wearable target. For the sake of simplicity in description, hereinafter, the detection device for the wearable target will be simply referred to as the detection device. The detection device can be implemented through a terminal or a server. The terminal is, for example, a mobile terminal, a fixed terminal or a portable terminal, such as a mobile phone, a multimedia computer, a multimedia tablet, a desktop computer, a notebook computer, a tablet computer, a device with a shooting function, etc. The server can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms, but is not limited thereto.
[0068] Some simple introductions will be made below to the application scenarios applicable to the technical solutions in the embodiments of the present application. It should be noted that the application scenarios introduced below are only used to illustrate the embodiments of the present application rather than to limit them. In the specific implementation process, the technical solutions provided in the embodiments of the present application can be flexibly applied according to actual needs.
[0069] Please refer toFigure 1 , which is a schematic diagram of an application scenario of a detection method for a wearable target provided by an embodiment of the present application. The application scenario schematic diagram includes a first image 110 and a detection device 120. The first image 110 can be collected by the detection device 120 itself. For example, if the detection device 120 is a device with a camera function, the first image 110 can also be sent to the detection device 120 after being collected by other camera devices.
[0070] Specifically, after the detection device 120 obtains the first image 110, it detects the first image 110 to determine whether the object to be detected in the first image 110 wears a target object. Among them, the specific process of detecting the wearable target object will be introduced in detail below.
[0071] As an embodiment, the target object in the embodiment of the present application can be a headdress, gloves, nameplate, employee ID card, etc. The headdress can be a headscarf, hat, etc. The object to be detected can be a person or an animal, and the present application does not make any limitations in this regard. For example, detecting whether a worker wears a safety helmet to ensure the safety of the worker, or detecting whether a pet wears a nameplate to avoid the pet getting lost, etc.
[0072] The detection method for wearable target objects provided by the embodiments of the present application can be applied to multiple scenarios, which will be introduced by examples below.
[0073] The first scenario: the entrance of a specific place.
[0074] A turnstile is set at the entrance of a specific place such as a kitchen or a workshop. A monitoring camera is installed at the turnstile. When each employee passes through the turnstile, the monitoring camera will take a picture of the employee, and detect whether the employee wears a target object according to the image. Only the employees who wear the target object can successfully pass through the turnstile.
[0075] The second scenario: the interior of a specific place.
[0076] Considering that after employees enter a specific place such as a kitchen or a workshop, they may also remove the target object for various reasons. Therefore, a monitoring camera is installed inside a specific place such as a kitchen or a workshop. The monitoring camera can capture the images of the employees in the place at any time, and detect whether the employees wear the target object according to the images, so as to report the situation of the employees who do not wear the target object in time.
[0077] The above introduces the application scenarios. Next, in combination with Figure 1 the application scenario shown, taking Figure 1 the detection device 120 in Figure 2 as an example to introduce the detection method for wearable target objects. Please refer to Figure 1 .
[0078] S201. Obtain a first image including the target part of the object to be detected.
[0079] Specifically, the detection device can collect a target image of the object to be detected in the target environment. The target environment refers to any place where it is necessary to detect whether the target object is worn, such as a kitchen, a workshop, etc. The target image can be sent to the detection device after being taken by other imaging devices. For example, other imaging devices are surveillance cameras in places such as kitchens and workshops. The target image can also be obtained by the detection device itself taking pictures. For example, the detection device is a device with a camera function.
[0080] Considering that the target image collected may be incomplete due to angle problems, after the detection device collects the target image, target detection algorithms such as FASTER-RCNN and YOLO can be used to detect whether there is a target part in the target image. The meaning of the object to be detected can be referred to the content discussed above and will not be elaborated here. The target part in the embodiments of the present application can be the head, hand, etc. Specifically, the target part can be a complete head or hand, such as the part above the neck, the part below the arm. The target part can also be an incomplete head or hand, such as the part above the nose or eyes, the part below the wrist. It should be noted that if the target part is the head, the target object can be a headdress. If the target part is the hand, the target object can be a glove.
[0081] Further, after the detection device detects whether there is a target part in the target image, according to different detection results, its processing process is different, which will be introduced separately below.
[0082] The first detection result is that there is no target part in the target image.
[0083] If there is no target part in the target image, it is impossible to further detect whether the target part is wearing the target object. In other words, the target image is invalid. Therefore, the detection device can delete the target image and re-collect the target image of the object to be detected in the target environment.
[0084] The second detection result is that there is a target part in the target image.
[0085] If the detection device determines that there is a target part in the target image, the target image can be directly used as the first image, or an image of the target part can be cropped from the target image based on a preset first rectangular frame as the first image. Specifically, after the detection device determines that there is a target part in the target image through the target detection algorithm, it can output the position information of the target part in the target image, that is, the position information of the first rectangular frame in the target image, and extract the pixel values of the area corresponding to the position information of the first rectangular frame from the target image, so as to obtain the first image.
[0086] In the embodiments of the present application, redundant image features are deleted, and subsequently, only the image of the area where the target part is located can be detected, which can improve the detection efficiency of the image.
[0087] Considering that the target image captured may be blurred due to reasons such as light and distance, therefore, in a possible embodiment, after the detection device acquires the target image of the target environment, it can detect the resolution of the target image. If it is determined that the resolution of the target image is less than the preset resolution, the target image can be deleted and the target image of the object to be detected in the target environment can be acquired again.
[0088] S202. Input the first image into the first model to obtain a first result.
[0089] Specifically, after the detection device obtains the first image, it can input the first image into the first model to obtain a first result. The first result is used to represent whether the target part wears the target object. The first model is obtained by training with a first sample set. The first sample set includes images marked for whether the target part of the historical object wears the target object. The historical object refers to a person or an animal in the training sample image. The first sample set specifically includes: images marked that the target part does not wear the target object, and images marked that the target part has worn the target object.
[0090] The first model can be various classification models, and the present application does not limit this. The values of the model parameters in the first model are obtained by training other devices and sent to the detection device, or are directly trained by the detection device. The process of the detection device training the first model is introduced by way of example below.
[0091] The detection device can use any sample image in the first sample set as the input of the first model. The error between the target category output by the first model and the actual category of the sample image is used as the feedback data of the first model. Through the feedback data, the values of the model parameters are continuously adjusted. After training with a large number of sample images, the model parameters of the first model are continuously updated so that the error between the target category determined by the first model and the actual category of the sample image is within the preset range, thereby obtaining the model parameters of the first model.
[0092] S203. In response to the first result indicating that the target part does not wear the target object, input the first image into the second model and the third model to obtain a second result and a third result.
[0093] After the detection device obtains the first result, if the first result indicates that the target part wears the target object, no subsequent processing is performed, and S204 is not executed. Considering that the detection result of the first model may be incorrect, in order to avoid false alarms, in response to the first result indicating that the target part does not wear the target object, the detection device inputs the first image into the second model and the third model to further verify the first result.
[0094] Specifically, the detection device inputs the first image into the second model to obtain a second result. The second result characterizes the attribute category of the target object in the first image. The second model is obtained by training with a second sample set, which includes images marked for target objects of different attribute categories. The different attribute categories include a first attribute category where the color of the target object is the target color, a second attribute category where the target part is not wearing the target object, and a third attribute category where the color of the target object is a non-target color. The target color refers to various colors that are difficult to detect, and the target color can be any one of black, white, or gray. Therefore, the second sample set specifically includes: images marked as the first attribute category, images marked as the second attribute category, and images marked as the third attribute category. The training process of the second model is similar to that of the first model. Please refer to the content described above, and it will not be elaborated here.
[0095] While inputting the first image into the second model, the detection device also inputs the first image into the third model to obtain a third result. The third result characterizes whether the first image is a difficult example image. In a difficult example image, the attribute category of the target object is the first attribute category. The third model is obtained by training with a third sample set, which includes images marked for whether the sample image is a difficult example image, specifically including: images marked as difficult examples and images marked as non-difficult examples. The training process of the third model is also similar to that of the first model. Please refer to the content described above, and it will not be elaborated here.
[0096] As an embodiment, the first model, the second model, and the third model can be three independent models, which are trained and obtained separately.
[0097] Alternatively, considering that training three models separately consumes a large amount of network resources, in a possible embodiment, the first model, the second model, and the third model are respectively three branches of a multi-task network. The three branches, that is, the three models, are obtained by jointly training with the multi-task network. It's just that when actually deployed, which branches to cut off is determined according to the different positions where they are used.
[0098] It should be noted that for a training sample image, it is not mandatory for it to have the dimensions of all three branches simultaneously. Only the dimensions that it does not have need to be marked as -1. When it is necessary to optimize the effect of a certain branch separately, for a sample image with only one dimension, the dimension of this sample image can be increased to improve the training effect without the need to add new sample images, so as to reduce the consumption of network resources. Additionally, the three branches share the backbone module, and their shared features contribute to the convergence of the backbone module. Moreover, the additional second and third branches help strengthen the network's overall recognition ability of the target object, and can explicitly provide local attention guidance to the network to highlight their different features, that is, to inform the network from which aspects to distinguish the features of the target object wearing the target color from those of the non-wearing target object, enabling the network to learn more fully the feature differences between the worn target object and the non-worn target object, thereby improving the classification effect.
[0099] As an embodiment, the first model is the first branch of the multi-task network, the second model is the second branch of the multi-task network, the third model is the third branch of the multi-task network, and the multi-task network further includes a backbone module. Among them, the backbone module conforms to the basic structure of a convolutional neural network, including but not limited to convolutional layers, non-linear layers, pooling layers, etc., such as inception, residual network (resnet), dense connection network (densnet), etc. The backbone module is mainly used for feature extraction. The processes of the multi-task network implementing S202 and S203 are introduced below.
[0100] The detection device inputs the first image into the backbone module, extracts the features of the first image to obtain the image features of the first image, and inputs the image features into the first branch to obtain a first result. In response to the first result indicating that the target part is not wearing the target object, the image features are input into the second branch to obtain a second result, and the image features are input into the third branch to obtain a third result.
[0101] Please refer to Figure 3 , which is a structural diagram of a multi-task network provided by an embodiment of the present application. The multi-task network includes a first branch 301, a second branch 302, a third branch 303, and a backbone module 304. The dotted arrows of the second branch 302 and the third branch 303 indicate that the second branch 302 and the third branch 303 may not exist. That is to say, in the actual use of the multi-task network, the second branch 302 and the third branch 303 may not be needed. For example, if the first result output by the first branch 301 indicates that the target part is wearing the target object, the second branch 302 and the third branch 303 are not executed.
[0102] Considering that the detection results of the target detection algorithm may be inaccurate, resulting in possible deviations in the first rectangular box. For example, there may be parts in the first rectangular box where there are no target objects. Therefore, in a possible embodiment, the first rectangular box is adjusted, and according to the obtained second rectangular box after adjustment, a second image is re-extracted from the target image and input into the second model and the third model for verification. The specific steps are as follows in S1.1 - S1.3.
[0103] S1.1. Perform N times of boundary jitter on the first rectangular box, and according to the position information of the first rectangular box in the target image, obtain the position information of N second rectangular boxes in the target image.
[0104] Among them, N is an odd number greater than or equal to 3. The main consideration is that 3 times of jitter can eliminate most of the deviations, and more jitter times will increase the time consumption instead. Therefore, N is generally taken as 3. Each boundary jitter means that the four boundaries of the first rectangular box are respectively expanded or contracted outward. In other words, the first rectangular box includes four boundaries: up, down, left, and right. When the first rectangular box performs boundary jitter each time, each of the four boundaries will jitter independently, that is, each boundary will be expanded or contracted outward. The jitter range of the left and right boundaries is based on the width of the first rectangular box as a reference, and the jitter range of the upper and lower boundaries is based on the height of the first rectangular box as a reference.
[0105] For example, if the width of the first rectangular box is w and the height is h, the jitter range of the left and right boundaries is (-0.1w, +0.1w), and the jitter range of the upper and lower boundaries is (-0.1h, +0.1h). The "-" indicates contraction inward, and the "+" indicates expansion outward. Considering that the upper margin of the rectangular box of the wearable target object is usually higher than that of the rectangular box of the non-wearable target object, therefore, the jitter range of the upper boundary can be larger than that of the lower boundary. Specifically, for example, the jitter range of the upper boundary is (-0.1h, +0.2h), and the jitter range of the lower boundary is (-0.1h, +0.1h).
[0106] It should be noted that among the N second rectangular boxes, there may be second rectangular boxes with the same size as the first rectangular box but different position information. For example, the upper boundary of the first rectangular box is expanded by 0.1h, the lower boundary is contracted by 0.1h, the left boundary is expanded by 0.1w, and the right boundary is contracted by 0.1w, which is equivalent to obtaining the second rectangular box after translating the first rectangular box upward and to the left.
[0107] S1.2. According to the position information of the N second rectangular boxes in the target image, extract N second images from the target image.
[0108] Specifically, after the detection device obtains the position information of N second rectangular frames, it can extract the pixel values of the area corresponding to the position information of each second rectangular frame from the target image, so as to obtain N second images.
[0109] S1.3. Input each second image into the second model and the third model to obtain a second result and a third result corresponding to each second image.
[0110] Specifically, after the detection device obtains N second images, it can input each second image into the second model to obtain a second result corresponding to each second image, and input each second image into the third model to obtain a third result corresponding to each second image. Among them, the meanings and training processes of the second model and the third model can refer to the content described above, which will not be elaborated here.
[0111] In the embodiment of the present application, multiple boundary jitters are performed on the first rectangular frame, which can eliminate the deviation problem of the first rectangular frame. Moreover, multiple verifications are performed on multiple second images, and multiple verifications can further improve the accuracy of the detection result of the wearable target. Considering that most of the staff are compliant in the wearable compliance scenario, that is, the number of people wearing the wearable target accounts for the majority of all people. When the first result indicates that the target part is not wearing the wearable target, boundary jitter is performed. If the first result indicates that the target part is wearing the wearable target, no subsequent processing will be performed, nor will boundary jitter be performed, which can save resource consumption.
[0112] S204. Determine whether the target part in the first image wears the wearable target based on the second result and the third result.
[0113] The detection device verifies the first result in different ways according to whether boundary jitter is performed on the first rectangular frame corresponding to the first image. The following will be introduced in different cases.
[0114] In the first case, the detection device does not perform boundary jitter on the first rectangular frame corresponding to the first image. In response to the first result indicating that the target part is not wearing the wearable target, the first image is directly input into the second model and the third model to obtain a second result and a third result corresponding to the first image.
[0115] For the first case, in response to the second result indicating that the attribute category of the wearable target in the first image is the second attribute category, and the third result indicating that the first image is not a difficult example image, it is determined that the target part in the first image is not wearing the wearable target. According to other situations of the second result and the third result, it can be determined that the target part in the first image is not wearing the wearable target. The following introduces the 5 situations included in other situations.
[0116] Case 1: The second result indicates that the attribute category of the target object in the first image is the first attribute category, and the third result indicates that the first image is a difficult example image.
[0117] Case 2: The second result indicates that the attribute category of the target object in the first image is the second attribute category, and the third result indicates that the first image is a difficult example image.
[0118] Case 3: The second result indicates that the attribute category of the target object in the first image is the third attribute category, and the third result indicates that the first image is a difficult example image.
[0119] Case 4: The second result indicates that the attribute category of the target object in the first image is the first attribute category, and the third result indicates that the first image is not a difficult example image.
[0120] Case 5: The second result indicates that the attribute category of the target object in the first image is the third attribute category, and the third result indicates that the first image is not a difficult example image.
[0121] Second case: After the detection device performs N boundary jitters on the first rectangular box corresponding to the first image to obtain N second images, each second image is input into the second model and the third model to obtain the second result and the third result corresponding to each second image.
[0122] For the second case, the detection device obtains the second results and the third results corresponding to the N second images, and cumulative voting can be performed. If the cumulative voting result meets the preset rule, it is determined that the target part in the first image has worn the target object.
[0123] Specifically, if the corresponding second result indicates that the attribute category of the target object in the first image is the second attribute category, and the number of second images whose corresponding third result indicates that the first image is not a difficult example image is greater than or equal to the preset value, it is determined that the target part in the first image has not worn the target object. If this number is less than the preset value, it is determined that the target part in the first image has not worn the target object. Among them, the preset value is proportional to the value of N. In other words, the preset value is proportional to the total number of multiple second images. For example, the preset value is β*N, and the value range of β can be [0.6, 0.75].
[0124] It should be noted that in order to ensure the accuracy of the detection result, the preset value can be the total number of multiple second images, that is, N. In other words, in response to each second result indicating that the attribute category of the target object in the first image is the second attribute category, and each third result indicating that the first image is not a difficult example image, it is determined that the target part in the first image has worn the target object.
[0125] As an embodiment, after the detection device determines that the target part in the first image is not wearing the target object, it can display an alarm message, which is used to indicate that the target part in the first image is not wearing the target object. The alarm message can be voice, text, image, etc. Taking the target object as a hat for example, the detection device can broadcast a voice message "Someone is not wearing a hat, please pay attention!", or the detection device has a display screen, and displays the text message "Someone is not wearing a hat, please pay attention!" on the display screen, and also displays the first image on the display screen, or the detection device can play a special ringtone to alarm for non-compliant wearing.
[0126] As an embodiment, after the detection device determines that the target part in the first image is not wearing the target object, it can determine the identity information of the object to be detected in the first image, and display the alarm message and the identity information of the object to be detected. The identity information is used to uniquely identify the object to be detected, such as the name of the object to be detected, the number of the object to be detected, etc. For the meaning of the alarm message, please refer to the content described above, and it will not be elaborated here.
[0127] Specifically, the detection device has pre-established an identity library, which includes the facial images of multiple objects to be detected and the identity information corresponding to each facial image. The detection device extracts the facial features of the object to be detected in the first image, and matches the facial features with the features of each facial image in the identity library one by one. If there is a successfully matched target facial image, it obtains the identity information corresponding to the target facial image from the identity library.
[0128] Taking the target object as a hat for example, if the detection device determines that the object to be detected corresponding to the first image is XXX, it can broadcast a voice message "XXX is not wearing a hat, please pay attention!", or the detection device has a display screen, and scrolls and displays the text message "XXX is not wearing a hat, please pay attention!" on the display screen, and also displays the first image on the display screen, or the detection device sends the text message "XXX is not wearing a hat" and the first image to the terminal of the management personnel.
[0129] For a clearer illustration of the detection method for wearing the target object, please refer to Figure 4 , which is a schematic flowchart of a detection method for wearing a target object provided by an embodiment of the present application Figure 2 Next, the detection method for wearing the target object will be introduced in detail in conjunction with Figure 4 .
[0130] The process starts. First, S401 is executed, that is, the first image is acquired.
[0131] S401. Acquire the first image.
[0132] For the manner in which the detection device acquires the first image and the meaning of the first image, please refer to the content described above, and it will not be elaborated here.
[0133] S402. Detect whether there is a target part.
[0134] The detection device detects whether there is a target part in the first image. If there is, S403 is executed, that is, it is detected whether the target object is worn on the target part. If not, S401 is continued to be executed, that is, the first image is acquired again. For how to detect the target part and the meaning of the target part, please refer to the content discussed above and will not be elaborated here.
[0135] S403. Detect whether the target object is worn on the target part.
[0136] The detection device can input the first image into the first model to obtain a first result. If the first result indicates that the target object has been worn on the target part, the process ends. If the first result indicates that the target object has not been worn on the target part, S404 is executed, that is, boundary jitter. For the training method of the first model, please refer to the content discussed above and will not be elaborated here.
[0137] S404. Boundary jitter.
[0138] The detection device performs boundary jitter on the first rectangular frame corresponding to the first image multiple times to obtain multiple second images. For the meaning of the first rectangular frame, the meaning of boundary jitter, and the process of how to obtain the second images, please refer to the content discussed above and will not be elaborated here.
[0139] S405. Detect the attribute category of the target object.
[0140] The detection device inputs each second image into the second model to obtain a second result, and obtains the attribute category of the target object according to the second result. For the training method of the second model, please refer to the content discussed above and will not be elaborated here.
[0141] S406. Detect whether it is an example image.
[0142] The detection device inputs each second image into the third model to obtain a third result, and obtains whether each second image is an example image according to the third result. For the training method of the third model, please refer to the content discussed above and will not be elaborated here.
[0143] S407. Cumulative voting.
[0144] The detection device performs cumulative voting according to the second results and third results corresponding to the multiple second images. For the specific voting method, please refer to the content discussed above and will not be elaborated here.
[0145] S408. Display an alarm message.
[0146] If the voting result meets the preset rules, the detection device displays an alarm message. For the meaning of the alarm message and the way of displaying the alarm message, please refer to the content described above, which will not be elaborated here.
[0147] After executing S408, the process ends.
[0148] Based on the same inventive concept, the present application also provides a detection device for wearable target objects. Please refer to Figure 5 , and the device includes:
[0149] An acquisition module 501, configured to acquire a first image of a target part including an object to be detected;
[0150] An obtaining module 502, configured to input the first image into a first model to obtain a first result, where the first result indicates whether the target part wears a target object; the first model is obtained by training with a first sample set, and the first sample set includes images marked with whether the target part of a historical object wears a target object;
[0151] The obtaining module 502 is further configured to, in response to the first result indicating that the target part does not wear a target object, input the first image into a second model and a third model to obtain a second result and a third result; wherein, the second result indicates the attribute category of the target object in the first image, and the third result indicates whether the first image is a difficult example image, and the attribute category of the target object in the difficult example image is a first attribute category; the second model is obtained by training with a second sample set, and the second sample set includes images marked with target objects of different attribute categories; the third model is obtained by training with a third sample set, and the third sample set includes images marked with whether the sample images are difficult example images;
[0152] A determination module 503, configured to determine whether the target part in the first image wears a target object based on the second result and the third result.
[0153] In a possible embodiment, the determination module 503 is specifically configured to:
[0154] In response to the second result indicating that the attribute category of the target object in the first image is a second attribute category, and the third result indicating that the first image is not a difficult example image, it is determined that the target part in the first image does not wear a target object.
[0155] In a possible embodiment, different attribute categories include a first attribute category where the color of the target object is a target color, a second attribute category where the target part does not wear a target object, and a third attribute category where the color of the target object is a non-target color.
[0156] In a possible embodiment, the target object is a headdress, and the target color is any one of black, white, or gray.
[0157] In a possible embodiment, the obtaining module 501 is specifically configured to:
[0158] Collect a target image of an object to be detected in a target environment;
[0159] If it is determined that there is a target part in the target image, then based on a preset first rectangular frame, extract an image of the target part from the target image as a first image.
[0160] In a possible embodiment, the obtaining module 502 is specifically configured to:
[0161] Perform N times of boundary jittering on the first rectangular frame, and obtain the position information of N second rectangular frames in the target image according to the position information of the first rectangular frame in the target image; where N is an odd number greater than or equal to 3, and each boundary jittering means expanding or contracting the four boundaries of the first rectangular frame respectively;
[0162] Extract N second images from the target image according to the position information of the N second rectangular frames in the target image;
[0163] Input each second image into a second model and a third model to obtain a second result and a third result corresponding to each second image.
[0164] In a possible embodiment, the determining module 503 is specifically configured to:
[0165] If the corresponding second result indicates that the attribute category of the target object in the first image is a second attribute category, and the number of second images corresponding to the third result indicating that the first image is not a difficult example image is greater than or equal to a preset value, then determine that the target part in the first image is not wearing the target object; where the preset value is proportional to the value of N.
[0166] In a possible embodiment, the first model is the first branch of a multi-task network, the second model is the second branch of the multi-task network, the third model is the third branch of the multi-task network, and the multi-task network further includes a backbone module;
[0167] The obtaining module 502 is specifically configured to:
[0168] Input the first image into the backbone module to obtain image features of the first image;
[0169] Input the image features into the first branch to obtain a first result;
[0170] In response to the first result indicating that the target part is not wearing the target object, input the image features into the second branch to obtain a second result, and input the image features into the third branch to obtain a third result.
[0171] As an embodiment, Figure 5 The described device can be used to executeFigure 2 and Figure 4 the detection method of the wearable target described in the embodiments shown, therefore, for the functions that each functional module of the device can achieve, etc., reference can be made to Figure 2 and Figure 4 the description of the embodiments shown, which will not be repeated here.
[0172] It should be noted that although several modules or sub-modules of the device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present invention, the features and functions of the two or more modules described above can be embodied in one module. Conversely, the features and functions of one module described above can be further divided and embodied by multiple modules.
[0173] Based on the same inventive concept, the embodiments of the present application also provide a detection device for a wearable target. This device is equivalent to the detection device discussed above. Please refer to Figure 6 , and this device includes:
[0174] At least one processor 601, and a memory 602 connected to at least one processor 601. In the embodiments of the present application, the specific connection medium between the processor 601 and the memory 602 is not limited. Figure 6 In Figure 6 it is taken as an example that the processor 601 and the memory 602 are connected through a bus 600. The bus 600 is represented by a thick line in Figure 6 , and the connection manners between other components are only for illustrative purposes and are not to be taken as limiting. The bus 600 can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation,
[0175] In the embodiments of the present application, the memory 602 stores instructions executable by at least one processor 601. By executing the instructions stored in the memory 602, at least one processor 601 can execute Figure 2 any of the detection methods of the wearable target described above. The processor 601 can also implement Figure 4 the functions of each module in the device shown.
[0176] Among them, the processor 601 is the control center of the device. It can connect various parts of the entire control device through various interfaces and lines. By running or executing the instructions stored in the memory 602 and calling the data stored in the memory 602, the various functions of the device and process data, so as to monitor the device as a whole.
[0177] In a possible design, the processor 601 may include one or more processing units. The processor 601 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communications. It can be understood that the above-mentioned modem processor may not be integrated into the processor 601 either. In some embodiments, the processor 601 and the memory 602 may be implemented on the same chip, and in some embodiments, they may also be separately implemented on independent chips.
[0178] The processor 601 may be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method for detecting a wearable target object disclosed in combination with the embodiments of the present application may be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.
[0179] As a non-volatile computer-readable storage medium, the memory 602 can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 602 may include at least one type of storage medium. For example, it may include flash memory, a hard disk, a multimedia card, a card-type memory, a random access memory (RAM), a static random access memory (SRAM), a programmable read-only memory (PROM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic memory, a magnetic disk, an optical disk, etc. The memory 602 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 602 in the embodiments of the present application may also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.
[0180] By programming the design of the processor 601, the code corresponding to the method for detecting a wearable target object introduced in the foregoing embodiments can be solidified into the chip, so that the chip can execute when running Figure 2 and Figure 4Steps of the detection method for the wearable target object shown. How to design and program the processor 601 is a well-known technology to those skilled in the art, and will not be elaborated here.
[0181] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer is caused to execute the detection method for the wearable target object as described in any of the foregoing discussions. Since the principle of solving problems by the above computer-readable storage medium is similar to that of the detection method for the wearable target object, the implementation of the above computer-readable storage medium can refer to the implementation of the method, and the repeated parts will not be elaborated.
[0182] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0183] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0184] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0185] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, so that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing steps for implementing the functions specified in one process or a plurality of processes and / or boxes Figure 1 one process or a plurality of processes and / or boxes Figure 1 in one box or a plurality of boxes.
[0186] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these modifications and variations.
Claims
1. A method for detecting a wearable target object, characterized in that, Including: Obtain a first image of a target part including the object to be detected, where the first image is an image of the target part cropped from a target image including the object to be detected based on a preset first rectangular frame; Input the first image into a first model to obtain a first result; the first result indicates whether the target part wears a target object, and the first model is obtained by training with a first sample set, where the first sample set includes images marked for whether the target parts of historical objects wear target objects; In response to the first result indicating that the target part does not wear a target object, input the first image into a second model and a third model to obtain a second result and a third result; where the second result indicates the attribute category of the target object in the first image, and the third result indicates whether the first image is a difficult example image, and the attribute category of the target object in the difficult example image is a first attribute category; the second model is obtained by training with a second sample set, where the second sample set includes images marked for target objects of different attribute categories; the third model is obtained by training with a third sample set, where the third sample set includes images marked for whether the sample images are difficult example images. Among them, inputting the first image into the second model and the third model to obtain the second result and the third result includes: performing N boundary jitters on the first rectangular frame corresponding to the first image to obtain N second images, and inputting each of the N second images into the second model and the third model to obtain the second result and the third result corresponding to each second image; Based on the second result and the third result, determine whether the target part in the first image wears a target object.
2. The method according to claim 1, characterized in that, Based on the second result and the third result, determining whether the target part in the first image wears a target object includes: In response to the second result indicating that the attribute category of the target object in the first image is a second attribute category and the third result indicating that the first image is not a difficult example image, determine that the target part in the first image does not wear a target object.
3. The method according to claim 2, characterized in that, The different attribute categories include a first attribute category where the color of the target object is a target color, a second attribute category where the target part does not wear the target object, and a third attribute category where the color of the target object is a non-target color.
4. The method according to claim 3, wherein The target object is a headdress, and the target color is any one of black, white, or gray.
5. The method according to any one of claims 1-4, characterized in that, Obtaining a first image of a target part including the object to be detected includes: Collect a target image of the object to be detected in a target environment; If it is determined that there is a target part in the target image, crop the image of the target part from the target image based on a preset first rectangular frame as the first image.
6. The method according to claim 5, wherein Performing N boundary jitters on the first rectangular frame corresponding to the first image to obtain N second images includes: Perform N times of boundary jittering on the first rectangular box, and obtain the position information of N second rectangular boxes in the target image according to the position information of the first rectangular box in the target image; each boundary jittering means expanding or shrinking the four boundaries of the first rectangular box respectively. Extract the N second images from the target image according to the position information of the N second rectangular boxes in the target image.
7. The method according to claim 6, characterized in that, Based on the second result and the third result, determine whether the target part in the first image wears the target object, including: If the corresponding second result indicates that the attribute category of the target object in the first image is the second attribute category, and the number of second images corresponding to the third result indicating that the first image is not a difficult example image is greater than or equal to a preset value, then determine that the target part in the first image does not wear the target object; wherein, the preset value is proportional to the value of N.
8. The method according to any one of claims 1 to 4, characterized in that, The first model is the first branch of the multi-task network, the second model is the second branch of the multi-task network, the third model is the third branch of the multi-task network, and the multi-task network further includes a backbone module. Input the first image into the first model to obtain a first result, including: Input the first image into the backbone module to obtain the image features of the first image. Input the image features into the first branch to obtain a first result. In response to the first result indicating that the target part does not wear the target object, input the first image into the second model and the third model to obtain a second result and a third result, including: In response to the first result indicating that the target part does not wear the target object, input the image features into the second branch to obtain a second result, and input the image features into the third branch to obtain a third result.
9. A detection device for a wearable target object, characterized in that, Include: An acquisition module, configured to acquire a first image including a target part of an object to be detected. An obtaining module, configured to input the first image into the first model to obtain a first result, where the first result indicates whether the target part wears the target object. The first model is obtained by training with a first sample set, and the first sample set includes images marked with whether the target part of a historical object wears the target object. The obtaining module is further configured to, in response to the first result indicating that the target object is not worn on the target part, input the first image into a second model and a third model to obtain a second result and a third result; wherein, the second result indicates the attribute category of the target object in the first image, the third result indicates whether the first image is a difficult example image, and the attribute category of the target object in the difficult example image is a first attribute category; the second model is obtained by training with a second sample set, and the second sample set includes images marked for target objects of different attribute categories; the third model is obtained by training with a third sample set, and the third sample set includes images marked for whether the sample images are difficult example images. Wherein, inputting the first image into the second model and the third model to obtain the second result and the third result includes: performing N times of boundary jitter on the first rectangular frame corresponding to the first image to obtain N second images, and inputting each of the N second images into the second model and the third model to obtain the second result and the third result corresponding to each second image; A determining module, configured to determine whether the target part in the first image wears the target object based on the second result and the third result.
10. A detection device for a wearable target object, characterized in that, Including: A memory, configured to store program instructions; A processor, configured to call the program instructions stored in the memory and execute the method according to any one of claims 1-8 according to the obtained program instructions.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program includes program instructions, and when the program instructions are executed by a computer, the computer is caused to execute the method according to any one of claims 1-8.
Citation Information
Patent Citations
Mask wearing standard degree detection method and device
CN113420675A
Wearing compliance detection method and device and computer readable storage medium
CN114220117A