A method and system for identifying the clothing of a target person
By simultaneously extracting shallow and deep features in the image recognition model, and combining cross entropy and binary cross entropy loss function training, the problem of insufficient overfitting and generalization ability of construction workers in the industrial environment is solved, and higher recognition accuracy and stronger generalization ability are achieved.
Patent Information
- Application Number
- CN202111112601.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-23
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-09-23
AI Technical Summary
The prior art has problems with insufficient overfitting and generalization capabilities in the dress recognition of construction workers in industrial environments, especially when there are fewer samples, deep learning networks cannot effectively utilize tooling color and texture features.
The image recognition model is used to extract shallow and deep features at the same time, and cross-entropy and binary cross-entropy loss functions are constructed respectively. Feature extraction and training are performed through the residual network ResNet50, and loss calculation is performed in combination with tooling area features to improve recognition accuracy and generalization capabilities.
Without increasing the calculation cost, the accuracy and generalization ability of target personnel's clothing recognition are significantly improved, avoiding the problem of shallow features being downplayed.
Smart Images

Figure CN113989733B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and particularly to a method and system for identifying the clothing of target personnel. Background Art
[0002] The statements in this part merely provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] In industrial scenarios, the clothing safety of construction workers is related to the safety risks of construction operations. When using intelligent algorithms to automatically identify and detect the clothing of construction workers, due to the limitations of limited training samples and non-uniform sample domains in convolutional neural networks, serious overfitting phenomena will occur when using convolutional neural networks for classification or detection networks. For example, Chinese Patent CN112149514A provides a method and system for detecting the safety clothing of construction workers, mainly aiming at the safety clothing management of workers in complex construction environments, and realizing automatic detection and warning functions by introducing deep learning technology, greatly improving the detection efficiency and detection accuracy; however, in the case of fewer industrial environment samples, it often leads to insufficient generalization ability. Chinese Patent CN110210338A discloses a method and system for detecting and identifying the clothing information of target personnel, which compares the detected personnel with those in the template library, and performs face recognition if there is face information, but it does not analyze the texture feature information of the clothing. The shallow features such as the color and texture of work clothes are important distinguishing features for recognition. However, in the deep features obtained after multiple downsamplings by the deep network, the above shallow features are diluted, resulting in insufficient recognition accuracy. Summary of the Invention
[0004] To solve the above problems, the present invention proposes a method and system for identifying the clothing of target personnel, which uses an image recognition model to simultaneously extract shallow features and deep features; and calculates the losses of the shallow features and deep features respectively, and constructs a total loss function of the model based on this, considering shallow features such as the color and texture features of work clothes in the personnel clothing images, and improving the recognition accuracy and generalization ability of the network.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] In a first aspect, the present invention provides a method for identifying the clothing of target personnel, including:
[0007] Obtain a set of personnel clothing images in a specific area and their corresponding work clothing area labels;
[0008] Generate an image mask label according to the work clothing area;
[0009] Extract features from the personnel clothing image with an image mask label using a pre - constructed image recognition model;
[0010] Preset a first down - sampling multiple and a second down - sampling multiple, construct loss functions for the features extracted based on the first down - sampling multiple and the second down - sampling multiple respectively, obtain the total loss function of the image recognition model based on the two types of loss functions, and train the image recognition model with this;
[0011] Obtain the work - clothing recognition result for the target personnel clothing image according to the trained image recognition model.
[0012] As an alternative implementation, the process of generating an image mask label according to the work - clothing area includes: if the corresponding position in the personnel clothing image belongs to the work - clothing area, the mask value is 1, otherwise it is 0.
[0013] As an alternative implementation, the first down - sampling multiple is greater than the second down - sampling multiple.
[0014] As an alternative implementation, the features extracted after sampling based on the first down - sampling multiple are deep features, and the features extracted after sampling based on the second down - sampling multiple are shallow features.
[0015] As an alternative implementation, the process of constructing a loss function for the features extracted based on the first down - sampling multiple includes: the deep feature f d , after passing through the classifier F c , the deep - feature representation logits is: logits = F c (f d ); construct a loss function according to the deep - feature representation and the cross - entropy loss.
[0016] As an alternative implementation, the process of constructing a loss function for the features extracted based on the second down - sampling multiple includes: the shallow feature f s , after passing through the classifier F m , the shallow - feature representation mask is: mask = F m (f s ); construct a loss function according to the shallow - feature representation and the binary cross - entropy loss.
[0017] As an alternative implementation, the total loss function of the image recognition model obtained based on the two types of loss functions is: loss = α * L c (logits, L)+β * L m (mask, M); where α and β are weight coefficients, L c is the cross - entropy loss, L mThe binary cross - entropy loss, Logits is the deep feature representation, and mask is the shallow feature representation.
[0018] In a second aspect, the present invention provides a target personnel clothing recognition system, including:
[0019] An acquisition module, configured to acquire a set of personnel clothing images within a specific area and their corresponding tooling area labels;
[0020] A mask generation module, configured to generate an image mask label according to the tooling area;
[0021] A feature extraction module, configured to extract features from the personnel clothing images with image mask labels by using a pre - constructed image recognition model;
[0022] A loss function construction module, configured to preset a first downsampling multiple and a second downsampling multiple, construct loss functions for the features extracted based on the first downsampling multiple and the second downsampling multiple respectively, obtain the total loss function of the image recognition model based on the two types of loss functions, and train the image recognition model with this;
[0023] An identification module, configured to obtain a tooling recognition result for the target personnel clothing image according to the trained image recognition model.
[0024] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in the first aspect is completed.
[0025] In a fourth aspect, the present invention provides a computer - readable storage medium for storing computer instructions. When the computer instructions are executed by the processor, the method described in the first aspect is completed.
[0026] Compared with the prior art, the beneficial effects of the present invention are:
[0027] In a target personnel clothing recognition method and system of the present invention, an image recognition model is constructed based on a deep learning network, and the image recognition model is used to extract shallow features and deep features simultaneously; this is because the shallow features such as the tooling color and texture features of the personnel clothing images are important distinguishing features. If the deep learning network is directly downsampled to obtain deep features, the above - mentioned shallow features will be diluted. Therefore, to avoid this problem, the present invention extracts shallow features and deep features simultaneously.
[0028] A target personnel clothing recognition method and system of the present invention thus uses a deep network to simultaneously extract shallow features and deep features, calculates the losses for the shallow features and deep features respectively, and constructs the total loss function of the model based on this, thereby avoiding overfitting.
[0029] A method and system for identifying the clothing of a target person according to the present invention utilize the deep network. Based on the basic architecture of the general classification network, the unique regional features of work clothing are used to learn shallow features, and organic fusion learning is performed based on deep features. Without any increase in computational cost, the accuracy and generalization ability of the classification network are significantly improved.
[0030] Advantages of additional aspects of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0032] Figure 1 It is a flowchart of the method for identifying the clothing of a target person provided in Embodiment 1 of the present invention;
[0033] Figure 2 It is a schematic diagram of feature extraction of the image recognition model provided in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0035] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0036] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0037] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0038] Embodiment 1
[0039] As Figure 1As shown in the figure, this embodiment provides a method for identifying the clothing of a target person, including:
[0040] S1: Obtain the set of images of the clothing of the people in a specific area and their corresponding tooling area labels;
[0041] S2: Generate an image mask label according to the tooling area;
[0042] S3: Extract features from the images of the clothing of the people with the image mask label by using a pre-constructed image recognition model;
[0043] S4: Preset the first downsampling factor and the second downsampling factor, construct loss functions for the features extracted based on the first downsampling factor and the second downsampling factor respectively, obtain the total loss function of the image recognition model based on the two types of loss functions, and train the image recognition model with this;
[0044] S5: Obtain the tooling recognition result for the image of the target person's clothing according to the trained image recognition model.
[0045] In this embodiment, an image recognition model is constructed based on a deep learning network, and the shallow features and deep features are extracted simultaneously by using the image recognition model; this is because the shallow features such as the tooling color and texture features of the images of the clothing of the people are important distinguishing features. If the deep features are directly obtained by downsampling using the deep learning network, the above-mentioned shallow features will be diluted. Therefore, in order to avoid this problem, this embodiment extracts the shallow features and deep features simultaneously.
[0046] Specifically:
[0047] In step S1, let the set of samples of the images of the clothing of the people in the specific area be T = {X0, X1,..., X n}, which contains n samples, and each picture is represented as X i , and the corresponding label is L i ; let the tooling area in each picture X i be represented as B i , where B i contains the tooling area as {xmin, ymin, xmax, ymax}.
[0048] In step S2, let the input picture be where h represents the height of the picture, w represents the width of the picture, and c represents the number of input channels of the picture;
[0049] Then, for the input picture X i generate the mask M i ;
[0050] If the corresponding position m in X i belongs to the corresponding tooling area B i, then its M i value corresponds to 1, otherwise 0; specifically as shown in the following formula:
[0051]
[0052] In step S3, set the feature extraction encoding F(x), then the feature extraction of the personnel dressing image with the image mask label is expressed as: f d , f s = F(X, M).
[0053] In this embodiment, as Figure 2 shown, preset the first downsampling multiple and the second downsampling multiple, and the first downsampling multiple is greater than the second downsampling multiple;
[0054] After sampling by the first downsampling multiple, deep features are extracted, and after sampling by the second downsampling multiple, shallow features are extracted;
[0055] That is, f d is the deep feature obtained after the image recognition model is sampled by the first downsampling multiple, and f s is the shallow feature obtained after the image recognition model is sampled by the first downsampling multiple;
[0056] Preferably, in this embodiment, f d is the deep feature obtained after the image recognition model is sampled by the downsampling multiple 32, and f s is the shallow feature obtained after the network is sampled by the downsampling multiple 8 or 16.
[0057] In step S4, the deep feature f d passes through the classifier F c to obtain the deep feature representation logits, specifically as shown in the following formula:
[0058] logits = F c (f d )
[0059] The shallow feature f s passes through the classifier F m to obtain the shallow feature representation mask, specifically as shown in the following formula:
[0060] mask = F m (f s )
[0061] Then, the total loss function of the image recognition model is expressed as:
[0062] loss = α * L c (logits, L) + β * L m (mask, M)
[0063] where α and β are weight coefficients;
[0064] L c represents the cross - entropy loss, and its specific formula is:
[0065]
[0066] In the formula, C is the number of classes; y ic is the sign function, which takes 1 if the true class of sample i is equal to c, and 0 otherwise; p ic is the predicted probability that the observed sample i belongs to class c.
[0067] L m represents the binary cross - entropy loss, and its specific formula is:
[0068]
[0069] In the formula, y is the true value and y′ is the estimated value.
[0070] Through the above loss function and deep - learning network, combined with sample data to train the image recognition model, a tooling image recognition model with higher accuracy and stronger generalization ability is obtained, and the dressed image of the target person is recognized according to the trained image recognition model.
[0071] In this embodiment, the image recognition model uses the ResNet50 residual network as the backbone network. In the ResNet50 network, features with downsampling factors of 16 and 32 are respectively extracted as shallow - layer features f s and deep - layer features f d ; then the corresponding encoders are respectively used for encoding to obtain the deep - layer feature representation logits and the shallow - layer feature representation mask; finally, the cross - entropy loss and the binary cross - entropy loss are used to optimize the entire classification network.
[0072] On the basis of the basic architecture of the general classification network in this embodiment, the unique regional features of the tooling clothing are used to learn the shallow - layer features and organically fuse and learn the deep - layer features, which significantly improves the accuracy and generalization ability of the classification network without any increase in computational cost.
[0073] Embodiment 2
[0074] This embodiment provides a dressed person recognition system for target persons, including:
[0075] An acquisition module, configured to acquire a set of dressed images of persons in a specific area and their corresponding tooling area labels;
[0076] A mask generation module, configured to generate an image mask label according to the tooling area;
[0077] A feature extraction module, configured to perform feature extraction on a person's dressed image with an image mask label by using a pre-constructed image recognition model;
[0078] A loss function construction module, configured to preset a first downsampling multiple and a second downsampling multiple, respectively construct loss functions for the features extracted based on the first downsampling multiple and the second downsampling multiple, obtain the total loss function of the image recognition model based on the two types of loss functions, and train the image recognition model with this;
[0079] An identification module, configured to obtain a workwear identification result for a target person's dressed image according to the trained image recognition model.
[0080] It should be noted here that the above modules correspond to the steps described in Embodiment 1. The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.
[0081] In more embodiments, there is also provided:
[0082] An electronic device, including a memory and a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in Embodiment 1 is completed. For the sake of brevity, it will not be elaborated here.
[0083] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0084] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random memory. For example, the memory may also store information about the device type.
[0085] A computer-readable storage medium, used to store computer instructions. When the computer instructions are executed by the processor, the method described in Embodiment 1 is completed.
[0086] The method in Embodiment 1 can be directly implemented by a hardware processor, or by a combination of hardware and software modules in the processor. The software module can be located in a mature storage medium in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0087] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in conjunction with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0088] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, they are not limitations on the protection scope of the present invention. Those skilled in the art should understand that various modifications or deformations that can be made without creative efforts by those skilled in the art on the basis of the technical solution of the present invention are still within the protection scope of the present invention.
Claims
1. A method for identifying the clothing of a target person, characterized in that, Including: Obtain a set of images of the clothing worn by people in a specific area and their corresponding work clothing area labels; Generate an image mask label according to the work clothing area, including: if the corresponding position in the image of the clothing worn by people belongs to the work clothing area, the mask value is 1, otherwise it is 0; Extract features from the image of the clothing worn by people with the image mask label using a pre-constructed image recognition model; Preset a first downsampling ratio and a second downsampling ratio, where the first downsampling ratio is greater than the second downsampling ratio; the features extracted after sampling based on the first downsampling ratio are deep features, and the features extracted after sampling based on the second downsampling ratio are shallow features, including the work clothing color and texture features of the image of the clothing worn by people; Construct loss functions for the deep features extracted based on the first downsampling ratio and the shallow features extracted based on the second downsampling ratio respectively, obtain the total loss function of the image recognition model based on the two types of loss functions, and train the image recognition model with this; Among them, use the unique regional features of the work clothing to learn the shallow features and conduct organic fusion learning of the deep features; Extract the features extracted when the second downsampling ratio is 8 or 16 as shallow features, and the features extracted when the first downsampling ratio is 32 as deep features, then encode them using the corresponding encoders respectively to obtain deep feature representations and shallow feature representations, and finally optimize the image recognition model using cross-entropy loss and binary cross-entropy loss; Obtain the work clothing recognition result for the target image of the clothing worn by people according to the trained image recognition model.
2. The method for identifying the clothing of a target person according to claim 1, wherein, The process of constructing a loss function based on the features extracted at the first downsampling multiple includes: the deep features obtained after sampling at the first downsampling multiple , after passing through the classifier , the deep feature representation logits are: ; construct a loss function based on the deep feature representation and cross-entropy loss.
3. The method for identifying the clothing of a target person according to claim 1, wherein The process of constructing a loss function based on the features extracted at the second downsampling multiple includes: the shallow features obtained after sampling based on the second downsampling multiple , after passing through the classifier , the mask of the shallow feature representation obtained is: ; construct a loss function according to the shallow feature representation and the binary cross-entropy loss.
4. The method for identifying the clothing of a target person according to claim 1, wherein The total loss function of the image recognition model obtained based on the two types of loss functions is: Among them, , are weight coefficients, is the cross-entropy loss, is the binary cross-entropy loss, Logits is the deep feature representation, and mask is the shallow feature representation.
5. A target person's clothing recognition system, characterized in that, Including: An acquisition module configured to obtain a set of images of the clothing worn by people in a specific area and their corresponding work clothing area labels; A mask generation module configured to generate an image mask label according to the work clothing area, including: if the corresponding position in the image of the clothing worn by people belongs to the work clothing area, the mask value is 1, otherwise it is 0; A feature extraction module configured to extract features from the image of the clothing worn by people with the image mask label using a pre-constructed image recognition model; A loss function construction module, configured to preset a first downsampling multiple and a second downsampling multiple, where the first downsampling multiple is greater than the second downsampling multiple; the features extracted after sampling based on the first downsampling multiple are deep features, and the features extracted after sampling based on the second downsampling multiple are shallow features, including the work clothing color and texture features of the personnel's clothing image; construct loss functions for the deep features extracted based on the first downsampling multiple and the shallow features extracted based on the second downsampling multiple respectively, obtain the total loss function of the image recognition model based on the two types of loss functions, and train the image recognition model with this; among them, use the unique regional features of the work clothing to learn the shallow features and perform organic fusion learning on the deep features; extract the features extracted when the second downsampling multiple is 8 or 16 as shallow features, and the features extracted when the first downsampling multiple is 32 as deep features, then encode them using the corresponding encoders respectively to obtain deep feature representations and shallow feature representations, and finally optimize the image recognition model using cross-entropy loss and binary cross-entropy loss; An identification module, configured to obtain a work clothing identification result for the target personnel's clothing image according to the trained image recognition model.
6. An electronic device, characterized in that, It includes a memory and a processor, as well as computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in any one of claims 1-4 is completed.
7. A computer-readable storage medium, characterized in that, For storing computer instructions, when the computer instructions are executed by the processor, the method described in any one of claims 1-4 is completed.
Citation Information
Patent Citations
Method and system for detecting and identifying dressing information of target person
CN110210338A
Safety dressing detection method and system for construction workers
CN112149514A
Single image defogging method based on context-guided generative adversarial network
CN112070688A
Worker dressing detection method and device, storage medium and electronic equipment
CN112598059A
Small target object image detection method and device, electronic equipment and storage medium
CN113255699A