Work uniform wearing recognition method and system
By using the CycleGAN model to process the new server images, multiple sets of old server images of different new and old levels were generated, and combined with these images, the Re-ID model was trained, which solved the problem that the work clothes recognition technology could not recognize samples of different new and old levels in construction site production scenarios, and improved the accuracy of model recognition.
Patent Information
- Application Number
- CN202210435538.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-24
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-04-24
AI Technical Summary
In construction site production scenarios, existing work clothes identification technology is difficult to effectively identify samples of different new and old levels of the same set of clothes, resulting in a decrease in the accuracy of the test results.
The CycleGAN model is used to process the new server images, generate multiple sets of old server images of different new and old degrees, and combine these images to train the Re-ID model to improve sample diversity and model recognition accuracy.
By generating multiple sets of old and old images of old and new clothes, the model's accuracy in identifying different old and new clothes of the same set of clothes is improved, and the problem of difficulty in sample collection is solved.
Smart Images

Figure CN114898400B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image recognition, and particularly to the technology of work uniform wearing recognition. Background Art
[0002] In the production scenario of a construction site, the safety of production is of crucial importance. Workers are required to wear work uniforms at the workplace, which can distinguish personnel identities and play a professional protective role for workers. In order to better manage the wearing situation of work uniforms, many work uniform recognition technologies have emerged. In a Chinese invention patent "Work Uniform Wearing State Detection Method, Device, Storage Medium and Electronic Device" with the publication number CN111860471B, aiming at the problem of real-time update of work uniform types, it is proposed to use reference feature vector clustering, and use the target feature vector to find the closest distance in the feature space and compare it with the threshold. However, in terms of sample collection in the article, only different angles of work uniforms are concerned, and the collection of samples with different degrees of newness and oldness is not considered. In the same factory area, it is very common for employees to wear the same type of clothes with different degrees of newness and oldness. Therefore, without considering the collection of samples with different degrees of newness and oldness, the accuracy of the detection results will be greatly reduced. To solve this problem, the conventional solution is to focus on collecting images of the same set of clothes with different degrees of newness during sample collection. However, it is very difficult to collect images of the same set of work uniforms (especially niche work uniforms) with different degrees of newness in a short period of time. Summary of the Invention
[0003] The purpose of this application is to provide a work uniform wearing recognition method and system, which proposes an effective sample enhancement method for the problem of difficult sample collection of the same set of clothes with different degrees of newness and oldness, increases sample diversity, and improves the accuracy of model recognition.
[0004] A Train a CycleGAN model using the new and old work uniform image sets;
[0005] B Collect multiple new work uniform images, and input each new work uniform image stacked with N groups of random noise into the generator of the trained CycleGAN model to output N groups of aged work uniform images corresponding to each new work uniform image;
[0006] C Train a Re-ID model using the multiple new work uniform images and their corresponding aged work uniform image sets;
[0007] D Obtain the image of the pedestrian target to be recognized;
[0008] E Use the trained Re-ID model to recognize the work uniform wearing of the pedestrian target image.
[0009] In a preferred example, the step C further includes the step of training a Re-ID model using the multiple new work uniform images and their corresponding aged work uniform image sets to obtain a feature vector extraction model;
[0010] Step E further includes the steps of: obtaining an image set of multiple work uniforms, calculating the feature vectors of each image of each work uniform according to the feature vector extraction model to construct a feature comparison library; calculating the clothing feature vector of the pedestrian target image according to the feature vector extraction model, and if the distance between the clothing feature vector and any feature vector in the feature comparison library is less than a first predetermined threshold, determining that the clothing in the pedestrian target image belongs to the corresponding work uniform.
[0011] In a preferred example, the method further includes:
[0012] If a new type of work uniform appears, calculate the feature vector of this type of work uniform according to the feature vector extraction model, add the feature vector corresponding to this type of work uniform to the feature comparison library, and set the corresponding threshold.
[0013] In a preferred example, step D further includes the following steps:
[0014] Obtain a sequence of video frame images from the video stream;
[0015] Perform moving object detection and pedestrian detection on the sequence of video frame images in turn to obtain a sequence of pedestrian images, and determine the first detected pedestrian as a candidate tracking target;
[0016] Continuously track the candidate tracking target, and the continuous tracking includes the following sub-steps (1) to (3): (1) At the first tracking, obtain multiple consecutive frames of images of the candidate tracking target from the sequence of pedestrian images, calculate the matching degree between its tracking position and detection position according to the multiple frames of images, and if the calculated matching degrees are all greater than the corresponding threshold, mark the corresponding images as the formal tracking state and construct a tracking queue; (2) Continue to track the images in the tracking queue. If the matching degree between the current tracking position and detection position of the images in the formal tracking state is not all greater than the corresponding threshold, modify it to the disappearing state; (3) Continue to track the images in the tracking queue. If the matching degree between the current tracking position and detection position of the images in the disappearing state is all greater than the corresponding threshold, modify it to the formal tracking state, otherwise modify it to the deleted state and delete the images in the deleted state from the tracking queue;
[0017] Select M images (M≥1) from the pedestrian images marked as the formal tracking state in the tracking queue as the pedestrian target images to be recognized.
[0018] In a preferred example, after continuously tracking the candidate tracking target, the following steps are further included:
[0019] Calculate the object box confidence of the pedestrian images calibrated as the formal tracking state in the tracking queue by using the object detection model;
[0020] If the object confidence of a pedestrian image is greater than the second predetermined threshold and the aspect ratio of the object box is greater than the third predetermined threshold, then retain the image; otherwise, delete the image from the queue.
[0021] In a preferred example, the method further includes the following steps:
[0022] If more than a predetermined percentage of the M selected pedestrian object images to be recognized are determined to be not wearing work uniforms, then determine that the pedestrian object is not wearing a work uniform.
[0023] In a preferred example, the multiple new uniform images include different types of new uniform images under specific conditions, and the image set of multiple work uniforms includes different types of work uniform images under the specific conditions;
[0024] The specific conditions include one or more of the following: different shooting angles, different behavioral actions, different shooting light conditions.
[0025] This application discloses a work uniform wearing recognition system, including:
[0026] A new uniform aging module, which includes a CycleGAN model. The new uniform aging module is used to pre-train the CycleGAN model with a new uniform and an old uniform image set, collect multiple new uniform images, input each new uniform image superimposed with N groups of random noises into the generator of the trained CycleGAN model, and output N groups of aged uniform images corresponding to each new uniform image;
[0027] An image acquisition module, which is used to obtain pedestrian object images to be recognized;
[0028] A recognition module, which includes a Re-ID model. The module is used to recognize the wearing of work uniforms in the pedestrian object images by using the trained Re-ID model, where the Re-ID model is trained by using the multiple new uniform images and their corresponding aged uniform image sets.
[0029] In a preferred example, the recognition module includes a feature vector extraction model, a feature comparison library, and a judgment module;
[0030] The feature vector extraction model is obtained by training a Re-ID model using the multiple new clothing images and their corresponding aged clothing image sets; the feature comparison library is constructed by obtaining image sets of multiple work uniforms and calculating the feature vectors of each image of each work uniform according to the feature vector extraction model; the judgment module is used to calculate the clothing feature vector of the pedestrian target image according to the feature vector extraction model, and if the distance between the clothing feature vector and any feature vector in the feature comparison library is less than a first predetermined threshold, it is determined that the clothing in the pedestrian target image belongs to the corresponding work uniform.
[0031] In a preferred example, the multiple new clothing images include new clothing images of different types under specific conditions, and the image sets of the multiple work uniforms include work uniform images of different types under the specific conditions;
[0032] The specific conditions include one or more of the following: different shooting angles, different behavioral actions, different shooting light conditions.
[0033] In the embodiments of the present application, there are at least the following advantages and beneficial effects:
[0034] First, the present invention uses the CycleGAN model to extract the image features of old clothes, so that the newly collected clothes can be aged through the model random noise, solving the problem of difficult collection of old clothes samples.
[0035] Second, compared with the addition of artificial features, the present invention collects samples under natural conditions for the model to learn abstract features, and then generates aged real samples, which will be closer to the real state.
[0036] Third, the feature extraction model obtained by training the Re-ID model using the image sets of clothes samples of different new and old degrees collected by the method of the present invention can recognize different new and old degrees of the same set of clothes as the same category, improving the recognition accuracy of the model.
[0037] Fourth, the present invention uses preprocessing of video foreground and background judgment to filter out invalid information, and only passes the frames with changes in the picture into the model for calculation, effectively reducing the operation energy consumption, improving the effective calculation rate, and saving calculation resources. In addition, the present invention uses the method of clothing Re-ID to train a convolutional neural network model for clothing feature extraction. Through the model trained by the present invention, in the mapped feature space, different clothes can be effectively distinguished without being affected by external conditions such as light and angle. As long as the work uniform comparison image library is updated, it can cope with the frequent change of work uniform types. At the same time, the present invention effectively reduces the false alarm rate through layer-by-layer filtering rules with low computational complexity while ensuring the real-time operation of the model.
[0038] The specification of this application records a large number of technical features, which are distributed in various technical solutions. If all possible combinations of technical features (i.e., technical solutions) of this application were to be listed, the specification would become overly lengthy. To avoid this problem, each technical feature disclosed in the above-mentioned invention content of this application, each technical feature disclosed in the following embodiments and examples, and each technical feature disclosed in the drawings can be freely combined with each other to form various new technical solutions (all of these technical solutions are considered to have been recorded in this specification), unless the combination of such technical features is technically infeasible. For example, in one example, features A + B + C are disclosed, and in another example, features A + B + D + E are disclosed, and features C and D are equivalent technical means that perform the same function. Technically, only one of them can be used, and it is impossible to use both simultaneously. Feature E can be combined with feature C technically. Then, the solution of A + B + C + D should not be considered to have been recorded due to technical infeasibility, while the solution of A + B + C + E should be considered to have been recorded. Brief Description of the Drawings
[0039] Figure 1 It is a schematic flowchart of a work uniform wearing recognition method according to the first embodiment of this application.
[0040] Figure 2 It is a schematic flowchart of a work uniform wearing recognition method according to an embodiment of this application.
[0041] Figure 3 It is a schematic structural diagram of a work uniform wearing recognition system according to the second embodiment of this application. Detailed Description of the Embodiments
[0042] In the following description, many technical details are presented for the reader to better understand this application. However, those of ordinary skill in the art can understand that even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in this application can still be implemented.
[0043] Term Explanation: The Re-ID model or algorithm described in this application is an algorithmic idea for domain problems. By training the model, the model can extract feature vectors that describe specific features of an image, and use the distance between feature vectors as a metric for classification. Based on such an algorithmic idea, we can use network models of different sizes and structures as the backbone network, such as ResNet, MobileNet, etc. The differences in the backbone network affect the accuracy and operation speed of the model.
[0044] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0045] The first embodiment of the present application relates to a method for identifying the wearing of work clothes, and its process is as shown in Figure 1 shown, and this method includes the following steps:
[0046] Step 101, training a CycleGAN model using a set of new clothes and old clothes images.
[0047] Step 102, collecting multiple new clothes images, and inputting each new clothes image superimposed with N groups of random noise into the generator of the trained CycleGAN model to output N groups of aged clothes images corresponding to each new clothes image.
[0048] Step 103, training a Re-ID model using the multiple new clothes images and their corresponding set of aged clothes images.
[0049] Step 104, obtaining an image of a pedestrian target to be recognized.
[0050] Step 105, using the trained Re-ID model to identify the wearing of work clothes in the pedestrian target image.
[0051] In the set of new clothes and old clothes images in step 101, the number of new and old clothes is preferably a large amount, and the old clothes are not limited to work clothes, and any kind of clothes for any purpose and type can be used. In step 101, a large number of old clothes samples are collected to train the CycleGAN model, so that the model can learn the image features of old clothes, and thus the generator of CycleGAN can be used to add random noise to age the new clothes samples.
[0052] Optionally, this step 103 further includes the step of: training a Re-ID model using the multiple new clothes images and their corresponding set of aged clothes images to obtain a feature vector extraction model.
[0053] Optionally, this step 105 further includes the steps of: obtaining a set of images of multiple work clothes, calculating the feature vectors of each image of each work clothes according to the feature vector extraction model to construct a feature comparison library; calculating the clothing feature vector of the pedestrian target image according to the feature vector extraction model, if the distance between the clothing feature vector and any feature vector in the feature comparison library is less than a first predetermined threshold, it is determined that the clothes in the pedestrian target image belong to the corresponding work clothes; otherwise, it is determined that no work clothes are worn, for example, an alarm prompt signal is issued.
[0054] Optionally, the images in the old work uniform image set in step 101, the new work uniform images in step 103, and the image set of multiple work uniforms in step 105 may include full-body images of upper and lower clothes, may also include separate upper and lower body clothing images, or may simultaneously include full-body images of upper and lower clothes, separate upper and lower body clothing images. For example, for separate upper and lower body clothing images, the upper and lower body clothes can be separately recognized in a targeted manner.
[0055] Optionally, the above-mentioned "multiple new work uniform images" may include new work uniform images of the same or different types under specific conditions, and the above-mentioned "image set of multiple work uniforms" may include work uniform images of the same or different types under the specific conditions.
[0056] Optionally, the specific conditions include, for example but not limited to, one or more of the following: different shooting angles, different action postures, different shooting light conditions.
[0057] Optionally, the method further includes the following steps: If a new type of work uniform appears, calculate the feature vector of this type of work uniform according to the feature vector extraction model, add the feature vector corresponding to this type of work uniform to the feature comparison library, and set the corresponding threshold.
[0058] Optionally, step 104 further includes the following sub-steps 1041 to 1043:
[0059] Step 1041: Obtain a sequence of video frame images from the video stream. Specifically, the data volume of a single high-definition picture is relatively large, and the frame rate of the camera for capturing pictures is also relatively high, capable of capturing 25 - 60 high-definition images per second. The data volume to be transmitted every day will be very large. Therefore, network cameras will encode the video frames before transmission, and at the server side, it is necessary to decode the received video frame encoded bitstream to obtain a single image. Optionally, the present invention uses the standard rtsp protocol to parse the bitstream to obtain video frame images.
[0060] Step 1042: Perform moving object detection and pedestrian detection on the sequence of video frame images in turn to obtain a sequence of pedestrian images, and determine the first detected pedestrian as the candidate tracking target.
[0061] When performing "moving object detection" in the above step 1042, the appearance of pedestrians in the video frame is a dynamic behavior, which must be caused by dynamic foregrounds in the video frame. Optionally, the present invention uses a background modeling method to model the background. Those that do not satisfy the model distribution must be foregrounds. Using this method, dynamic targets can be detected. For the background model, we model each pixel point, establish multiple Gaussians for each pixel point, and continuously update the mean and variance of the Gaussians on the time axis. When the pixel points corresponding to the current frame satisfy the distribution of these multiple Gaussians, they are background points; otherwise, they are foreground points. This algorithm has a low computational complexity and can filter out most of the image frames without moving objects, thus reducing a large amount of computational work. Throughout the 24 hours of a day, there are only very few moments when there are people passing by the camera in scenes with dynamic targets, and most of the invalid video frame calculations can be filtered out. The "pedestrian detection" in the above step 1042 is to detect pedestrians in the moving video frames. Optionally, the present invention uses a convolutional neural network to design an effective network structure and loss function, and obtains effective convolutional weights through gradient descent learning. Optionally, this application uses yolov5 for pedestrian detection. Yolov5 is just one of many object detection algorithms and is an anchor based algorithm, which is currently the best in this type of algorithm. In addition, there is also a type of anchor free algorithm that can complete the object detection task. When the order of magnitude of the number of objects is different, the operation speeds of the two methods have their own advantages and disadvantages. In the scenarios addressed by the present invention, when the pedestrians on the construction site are not dense, the anchor based algorithm is more applicable.
[0062] Step 1043: Continuously track the candidate tracking target. This continuous tracking includes the following sub-steps (1) to (3): (1) At the first tracking, obtain consecutive multiple frames of images of the candidate tracking target from the pedestrian image sequence, calculate the matching degree between its tracking position and detection position based on these multiple frames of images. If the calculated matching degrees are all greater than the corresponding thresholds, mark the corresponding images as the formal tracking state and construct a tracking queue; (2) Continue to track the images in the tracking queue. If the matching degree between the current tracking position and detection position of the images in the formal tracking state is not all greater than the corresponding thresholds, modify it to the disappearing state; (3) Continue to track the images in the tracking queue. If the matching degree between the current tracking position and detection position of the images in the disappearing state is all greater than the corresponding thresholds, modify it to the formal tracking state; otherwise, modify it to the deleted state and delete the images in the deleted state from the tracking queue. This step can avoid the occurrence of repeated alarms.
[0063] Specifically, for the "pedestrian target tracking" in step 1043 above, a single-object tracker optionally selects the Kalman tracking algorithm. When there are multiple targets to be tracked in the video, they are managed using an array. Each target is a tracker, and depending on the detection situation, the tracker has different states: 1. When a target is first detected, it may be a real target or a false detection. It is necessary to continuously examine the matching situation between the tracking positions and the detection positions in several consecutive frames. In the present invention, this state of the tracking target is defined as a candidate tracking target; 2. After several consecutive frames of observation, if there is a relatively high matching degree between the tracking position and the detection position, it can be considered that the current target is a reliable tracking target, and the state of this target is changed to a formal tracking target; 3. When a formal tracking target is approaching disappearance or the detection is unstable due to reasons such as lighting and occlusion, there will be a mismatch between the detection position and the tracking position. In this case, we change the state of this target to a disappearing state. During the subsequent tracking process, if it can be rematched again, it will be changed to a formal tracking target; if the detection box still cannot be matched in multiple subsequent frames, this target will be deleted from the tracking queue; 4. The state of the deleted target is the deleted state, which is the target that finally disappears from the screen.
[0064] Step 1044: Select M images from the pedestrian images marked as the formal tracking state in the tracking queue, where M≥1, as the pedestrian target images to be recognized.
[0065] Optionally, after step 1043, the following steps ① and ② are further included: ① Calculate the confidence of the target box of the pedestrian images marked as the formal tracking state in the tracking queue using the target detection model; ② If the target confidence of a pedestrian image is greater than the second predetermined threshold and the aspect ratio of the target box is greater than the third predetermined threshold, then retain this image, otherwise delete this image from the queue. Steps ① and ② mainly address the problem that if a pedestrian is largely occluded, it is likely to cause deviation in the extraction of work uniform features. In the prior art, keypoint detection is generally used. However, if keypoint detection is used, there will be as many as three deep neural network models in the entire pipeline, namely the target detection model, the keypoint detection model, and the work uniform feature extraction model, and the operation speed will be very limited. Therefore, the present invention optionally uses some judgment rules with low computational complexity to preliminarily filter the occlusion situation: First, it is the confidence of the target box output by the target detection model. If the confidence is higher than the threshold, then enter the next judgment rule, otherwise directly skip and do not perform feature extraction on this pedestrian target in this frame. Second, it is to judge the aspect ratio of the target box. If the ratio of "height / width" is greater than 2, it is considered that this is a pedestrian in a reasonable upright state, and enter step 6 for feature extraction, otherwise directly skip and return to the step of obtaining image frames from the video stream.
[0066] Optionally, after step 1044, the following steps are further included: If more than a predetermined percentage of the selected M pedestrian target images to be recognized are determined to be without work uniforms, it is determined that the pedestrian target is not wearing a work uniform. It should be noted that although the above steps ① and ② filter upright pedestrians through rules, there will still be situations where filtering fails. Considering that pedestrians are moving targets, misidentifications often occur only in a few frames. Therefore, in the present invention, a time window is used for smoothing filtering here. When the number of frames with work uniform non-wearing alarms in the time window is greater than the set percentage threshold, an alarm signal is given. In this way, the false alarm rate can be effectively reduced with almost no increase in computational complexity.
[0067] It should be noted that in this application, the "new uniform" and "old uniform" are abbreviations for new clothes and old clothes respectively.
[0068] The second embodiment of this application relates to a work uniform wearing recognition system, the structure of which is as Figure 3 shown. The work uniform wearing recognition system includes a new uniform aging module, an image acquisition module, and a recognition module.
[0069] The new uniform aging module includes a CycleGAN model. The new uniform aging module is used to pre-train the CycleGAN model using a set of new uniform and old uniform images, collect multiple new uniform images, input each new uniform image superimposed with N groups of random noise into the generator of the trained CycleGAN model, and output N groups of aged uniform images corresponding to each new uniform image.
[0070] The image acquisition module is used to obtain pedestrian target images to be recognized.
[0071] The recognition module includes a Re-ID model. The module is used to recognize the work uniform wearing in the pedestrian target image using the trained Re-ID model, where the Re-ID model is trained using the set of multiple new uniform images and their corresponding aged uniform images.
[0072] Optionally, the recognition module includes a feature vector extraction model, a feature comparison library, and a judgment module. The feature vector extraction model is obtained by training the Re-ID model using the set of multiple new uniform images and their corresponding aged uniform images; the feature comparison library is constructed by obtaining a set of images of various work uniforms and calculating the feature vectors of each image of each work uniform according to the feature vector extraction model; the judgment module is used to calculate the clothing feature vector of the pedestrian target image according to the feature vector extraction model. If the distance between the clothing feature vector and any feature vector in the feature comparison library is less than the first predetermined threshold, it is determined that the clothing in the pedestrian target image belongs to the corresponding work uniform.
[0073] Optionally, the multiple new work uniform images include different types of new work uniform images under specific conditions, and the set of images of the multiple work uniforms includes different types of work uniform images under the specific conditions.
[0074] Optionally, the specific conditions include one or more of the following: different shooting angles, different behavioral actions, and different shooting light conditions.
[0075] The first implementation manner is a method implementation manner corresponding to this implementation manner. The technical details in the first implementation manner can be applied to this implementation manner, and the technical details in this implementation manner can also be applied to the first implementation manner.
[0076] It should be noted that those skilled in the art should understand that the implementation functions of the various modules shown in the above implementation manners of the work uniform wearing recognition system can be understood with reference to the relevant descriptions of the foregoing work uniform wearing recognition method. The functions of the various modules shown in the above implementation manners of the work uniform wearing recognition system can be implemented by a program (executable instruction) running on a processor, or can also be implemented by specific logic circuits. If the work uniform wearing recognition system in the embodiments of the present application is implemented in the form of software function modules and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), magnetic disks, or optical discs that can store program codes. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.
[0077] Accordingly, an embodiment of the present application further provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the method embodiments of the present application. The computer-readable storage medium includes permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of the computer storage medium include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. As defined herein, the computer-readable storage medium does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0078] In addition, an embodiment of the present application further provides a work uniform wearing recognition system, which includes a memory for storing computer-executable instructions and a processor; the processor is configured to implement the steps in the above method embodiments when executing the computer-executable instructions in the memory. Wherein, the processor may be a central processing unit (Central Processing Unit, abbreviated as "CPU"), or may also be other general-purpose processors, digital signal processors (Digital Signal Processor, abbreviated as "DSP"), application specific integrated circuits (Application Specific Integrated Circuit, abbreviated as "ASIC"), etc. The aforementioned memory may be a read-only memory (abbreviated as "ROM"), a random access memory (abbreviated as "RAM"), a flash memory, a hard disk or a solid state drive, etc. The steps of the methods disclosed in the embodiments of the present invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules in the processor.
[0079] It should be noted that in the application documents of this patent, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one" does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element. In the application documents of this patent, if it is mentioned that an act is performed according to a certain element, it means that the act is performed at least according to that element, including two cases: the act is performed only according to that element, and the act is performed according to that element and other elements. Expressions such as multiple, many times, various, etc. include 2, 2 times, 2 kinds, as well as more than 2, more than 2 times, more than 2 kinds.
[0080] All documents mentioned in this application are considered to be integrally included in the disclosure content of this application so that they can be used as a basis for modification when necessary. In addition, it should be understood that the above are only preferred embodiments of this specification and are not used to limit the protection scope of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of one or more embodiments of this specification shall be included within the protection scope of one or more embodiments of this specification.
Claims
1. A method for identifying the wearing of work uniforms, characterized in that, it includes: Step A: Train a CycleGAN model using a set of new and old uniform images; Step B: Collect multiple new uniform images, and input each new uniform image superimposed with N groups of random noise into the generator of the trained CycleGAN model to output N groups of aged uniform images corresponding to each new uniform image; Step C: Train a Re-ID model using the set of multiple new uniform images and their corresponding aged uniform images; Step D: Obtain an image of a pedestrian target to be identified; Step E: Use the trained Re-ID model to identify the wearing of work uniforms in the pedestrian target image.
2. The method for identifying the wearing of work uniforms according to claim 1, characterized in that, Step C further includes the steps of: training a Re-ID model using the set of multiple new uniform images and their corresponding aged uniform images to obtain a feature vector extraction model; Step E further includes the steps of: obtaining a set of images of multiple work uniforms, calculating the feature vectors of each image of each work uniform according to the feature vector extraction model to construct a feature comparison library; calculating the clothing feature vector of the pedestrian target image according to the feature vector extraction model, if the distance between the clothing feature vector and any feature vector in the feature comparison library is less than a first predetermined threshold, it is determined that the clothing in the pedestrian target image belongs to the corresponding work uniform.
3. The method for identifying the wearing of work uniforms according to claim 2, characterized in that, The method further includes: If a new type of work uniform appears, calculate the feature vector of this type of work uniform according to the feature vector extraction model, add the feature vector corresponding to this type of work uniform to the feature comparison library, and set the corresponding threshold.
4. The method for identifying the wearing of work uniforms according to claim 1, characterized in that, Step D further includes the following steps: Obtain a sequence of video frame images from a video stream; Perform moving object detection and pedestrian detection on the sequence of video frame images in turn to obtain a sequence of pedestrian images, and determine the first detected pedestrian as a candidate tracking target; Perform continuous tracking on the candidate tracking target, and the continuous tracking includes the following sub-steps (1) to (3): (1) At the first tracking, obtain multiple consecutive frames of images of the candidate tracking target from the sequence of pedestrian images, calculate the matching degree between its tracking position and detection position according to the multiple frames of images, if the calculated matching degrees are all greater than the corresponding threshold, mark the corresponding image as the formal tracking state and construct a tracking queue; (2) Continue to track the images in the tracking queue, if the matching degree between the current tracking position and detection position of the image in the formal tracking state is not all greater than the corresponding threshold, modify it to the disappearing state; (3) Continue to track the images in the tracking queue, if the matching degree between the current tracking position and detection position of the image in the disappearing state is all greater than the corresponding threshold, modify it to the formal tracking state, otherwise modify it to the deleted state and delete the image in the deleted state from the tracking queue; Select M images from the pedestrian images marked as the formal tracking state in the tracking queue as the images of the pedestrian target to be identified, M≥1.
5. The work uniform wearing recognition method according to claim 4, characterized in that, after continuously tracking the candidate tracking target, the following steps are further included: Using an object detection model to calculate the object box confidence of the pedestrian images calibrated as the formal tracking state in the tracking queue; If the object confidence of a pedestrian image is greater than the second predetermined threshold and the aspect ratio of the object box is greater than the third predetermined threshold, then retain the image, otherwise delete the image from the queue.
6. The work uniform wearing recognition method according to claim 4 or 5, characterized in that, the method further includes the following steps: If more than a predetermined percentage of the M selected pedestrian target images to be recognized are determined to be without work uniforms, then it is determined that the pedestrian target is not wearing a work uniform.
7. The work uniform wearing recognition method according to claim 2, characterized in that, the multiple new uniform images include different types of new uniform images under specific conditions, and the image set of the multiple work uniforms includes different types of work uniform images under the specific conditions; the specific conditions include one or more of the following: different shooting angles, different behavioral actions, different shooting light conditions.
8. A work uniform wearing recognition system, characterized in that, including: A new uniform aging module, which includes a CycleGAN model. The new uniform aging module is used to pre-train the CycleGAN model with a new uniform and an old uniform image set, collect multiple new uniform images, input each new uniform image stacked with N groups of random noise into the generator of the trained CycleGAN model, and output N groups of aged uniform images corresponding to each new uniform image; An image acquisition module, which is used to acquire pedestrian target images to be recognized; A recognition module, which includes a Re-ID model. The module is used to recognize the wearing of work uniforms in the pedestrian target images by using the trained Re-ID model, wherein the Re-ID model is trained with the multiple new uniform images and their corresponding aged uniform image sets.
9. The work uniform wearing recognition system according to claim 8, characterized in that, the recognition module includes a feature vector extraction model, a feature comparison library, and a judgment module; The feature vector extraction model is obtained by training the Re-ID model with the multiple new uniform images and their corresponding aged uniform image sets; the feature comparison library is constructed by obtaining an image set of multiple work uniforms and calculating the feature vectors of each image of each work uniform according to the feature vector extraction model; the judgment module is used to calculate the clothing feature vector of the pedestrian target image according to the feature vector extraction model. If the distance between the clothing feature vector and any feature vector in the feature comparison library is less than the first predetermined threshold, then it is determined that the clothing in the pedestrian target image belongs to the corresponding work uniform.
10. The work uniform wearing recognition system according to claim 9, characterized in that, the multiple new uniform images include different types of new uniform images under specific conditions, and the image set of the multiple work uniforms includes different types of work uniform images under the specific conditions; the specific conditions include one or more of the following: different shooting angles, different behavioral actions, different shooting light conditions.
Citation Information
Patent Citations
A method and system for identifying work clothes based on feature retrieval
CN111860471B
Feature retrieval-based work clothes wearing identification method and system
CN111860471A
KR20210069354A