Training method of wearing detection model, wearing detection method and related equipment
By training the feature extraction, attribute prediction, and category prediction layers of the wear detection model, the problem of low detection accuracy caused by high clothing similarity is solved, and higher wear detection accuracy is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG DAHUA TECH CO LTD
- Filing Date
- 2022-07-27
- Publication Date
- 2026-04-21
AI Technical Summary
In existing technologies, due to the wide variety and high similarity of clothing types, the accuracy of wear detection is low, making it difficult to accurately distinguish similar clothing.
A wearable detection model training method is adopted, including a feature extraction layer, an attribute prediction layer, and a category prediction layer. Sample extraction features, wearable attribute features, and identity features are obtained through training sample images. The training loss is calculated and the model parameters are adjusted to improve detection accuracy.
By training the model from multiple aspects, including clothing attributes, identity categories, and image features, the detection accuracy of the clothing detection model has been significantly improved, and confusion in the identification of similar clothing has been reduced.
Smart Images

Figure CN115424294B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of detection technology, specifically to a training method for a wearable detection model, a wearable detection method, and related equipment. Background Technology
[0002] Currently, feature comparison methods are used to determine whether a target object is wearing the target clothing. However, in practical applications of clothing detection, due to the wide variety of clothing types and the high similarity between clothing of the same or similar colors and styles, similar clothing is easily confused during the clothing detection process, resulting in low accuracy of clothing detection. Summary of the Invention
[0003] This application provides a training method for a wear detection model, a wear detection method, and related equipment, which can improve the wear detection accuracy of the trained wear detection model.
[0004] To address the aforementioned technical problems, the technical solution adopted in this application is as follows: A training method for a wearable detection model is provided. The wearable detection model includes a feature extraction layer, an attribute prediction layer, and a category prediction layer. The training method includes: acquiring training sample images; inputting the training sample images into the feature extraction layer to obtain sample extracted features; inputting the sample extracted features into the attribute prediction layer to obtain sample wearable attribute features; inputting the sample extracted features into the category prediction layer to obtain sample identity features; calculating the training loss corresponding to the training sample images based on the sample wearable attribute features, sample identity features, and sample extracted features; and training the wearable detection model based on the training loss to obtain the trained wearable detection model.
[0005] To address the aforementioned technical problems, another technical solution adopted in this application is: providing a wearable detection method, which includes: acquiring an image to be identified and a set of comparison images; the set of comparison images includes multiple comparison images and wearable attribute labels and identity labels corresponding to each comparison image; inputting the image to be identified into a trained wearable detection model to obtain the features to be identified corresponding to the image to be identified, wherein the trained wearable detection model is trained using the training method of the wearable detection model in the above technical solution; inputting the comparison images into the trained wearable detection model to obtain the comparison features corresponding to the comparison images; and obtaining a wearable detection result based on the features to be identified, the comparison features, and the wearable attribute labels; the wearable detection result includes whether the wear of the target object in the image to be identified meets the wearable requirements corresponding to the identity label.
[0006] To solve the above-mentioned technical problems, another technical solution adopted in this application is: to provide a model training device, which includes a memory and a processor connected to each other, wherein the memory is used to store a computer program, and when the computer program is executed by the processor, it is used to implement the training method of the wearable detection model in the above-mentioned technical solution.
[0007] To solve the above-mentioned technical problems, another technical solution adopted in this application is: to provide a wearable detection device, which includes a memory and a processor connected to each other, wherein the memory is used to store a computer program, and when the computer program is executed by the processor, it is used to implement the wearable detection method in the above-mentioned technical solution.
[0008] To solve the above-mentioned technical problems, another technical solution adopted in this application is to provide a computer-readable storage medium for storing a computer program, which, when executed by a processor, is used to implement the training method or wear detection method of the wear detection model in the above-mentioned technical solution.
[0009] The beneficial effects of this application through the above scheme are as follows: by inputting training sample images into the feature extraction layer, sample extraction features are obtained; then, the sample extraction features are input into the attribute prediction layer and the category prediction layer respectively to obtain sample wearing attribute features and sample identity features; then, based on the sample wearing attribute features, sample identity features, and sample extraction features, the training loss corresponding to the training sample images is calculated, and the training loss is used to train the wear detection model to obtain the trained wear detection model; by training the wear detection model from three aspects—wearing attributes, identity category, and image features—the training effect of the wear detection model is greatly improved, thereby improving the wear detection accuracy of the trained wear detection model. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:
[0011] Figure 1 This is a flowchart illustrating an embodiment of the training method for the wearable detection model provided in this application;
[0012] Figure 2 This is a flowchart illustrating another embodiment of the training method for the wearable detection model provided in this application;
[0013] Figure 3 This is a schematic diagram of the structure of an embodiment of the wearable detection model provided in this application;
[0014] Figure 4 This is a schematic flowchart of an embodiment of the wearable detection method provided in this application;
[0015] Figure 5 This is a schematic diagram of another embodiment of the wearable detection model provided in this application;
[0016] Figure 6 yes Figure 4 The flowchart of step 44 in the illustrated embodiment is shown.
[0017] Figure 7 This is a schematic diagram of the structure of an embodiment of the model training device provided in this application;
[0018] Figure 8 This is a schematic diagram of the structure of an embodiment of the wearable detection device provided in this application;
[0019] Figure 9 This is a schematic diagram of an embodiment of the computer-readable storage medium provided in this application. Detailed Implementation
[0020] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be particularly noted that the following embodiments are for illustrative purposes only and do not limit the scope of the application. Similarly, the following embodiments are only some, not all, embodiments of the present application, and all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of the present application.
[0021] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0022] It should be noted that the terms "first," "second," and "third" in this application are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first," "second," or "third" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0023] Please see Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the training method for the wearable detection model provided in this application. The wearable detection model may include a feature extraction layer, an attribute prediction layer, and a category prediction layer. The method includes:
[0024] Step 11: Obtain training sample images.
[0025] The training sample images are images that include the target object, which can be a person or other animal. The clothing detection model is used to detect whether the clothing of the target object in the image meets the wearing conditions.
[0026] Step 12: Input the training sample images into the feature extraction layer to obtain the sample extracted features.
[0027] Training sample images can be input into the feature extraction layer to obtain sample extracted features, which may include image features corresponding to the training sample images. Specifically, the feature extraction layer can be used to extract features from the training sample images to obtain feature maps corresponding to the training sample images. Then, the feature maps are reconstructed and feature embedding processes are performed to adjust the dimensions of the feature maps, thereby obtaining sample extracted features in the target dimension. Understandably, the feature extraction layer can be a network structure used for feature extraction in this technical field, such as a Convolutional Neural Network (CNN), which is not limited here.
[0028] Step 13: Input the extracted features of the sample into the attribute prediction layer to obtain the sample's wearing attribute features.
[0029] The sample wearing attribute features can include the wearing attributes of the target object in the training sample image. Wearing attributes can include clothing color or clothing style, etc. The attribute prediction layer can be used to predict the wearing attributes based on the sample extracted features, thereby obtaining the sample wearing attribute features.
[0030] Step 14: Input the extracted features of the sample into the category prediction layer to obtain the sample identity features.
[0031] The sample identity features may include the clothing category worn by the target object, which may be related to the target object's identity / job position, such as firefighter uniform or chef's uniform. It can be understood that the wear detection model in this embodiment can be applied to the detection of work clothes, and the corresponding clothing category can be various work clothes categories. In other embodiments, the wear detection model can also be used to detect other non-work clothes, such as school uniforms or stage costumes, etc., which are not limited here.
[0032] Step 15: Calculate the training loss corresponding to the training sample image based on the sample wearing attribute features, sample identity features, and sample extraction features.
[0033] After obtaining the sample clothing attribute features, sample identity features, and sample extraction features of the training sample images, the training loss corresponding to the training sample images can be calculated based on the sample clothing attribute features, sample identity features, and sample extraction features.
[0034] Step 16: Train the wear detection model based on the training loss to obtain the trained wear detection model.
[0035] The parameters of the wearable detection model are adjusted based on the training loss. It is then determined whether the adjusted model meets the training termination condition. If yes, the trained wearable detection model is obtained; otherwise, the process returns to the step of obtaining training sample images, i.e., returning to step 11, until the adjusted model meets the training termination condition. Specifically, the training termination condition may include: loss convergence, i.e., the difference between the previous loss and the current loss value is less than a set value; determining whether the current loss value is less than a preset loss, which is a pre-set loss threshold; the number of training iterations reaching a set value (e.g., 10,000 training iterations); or the accuracy obtained when testing with a test set reaching a set condition (e.g., exceeding a preset accuracy).
[0036] This embodiment obtains sample extracted features by inputting training sample images into the feature extraction layer; then, the sample extracted features are input into the attribute prediction layer and the category prediction layer respectively to obtain sample wearing attribute features and sample identity features; then, based on the sample wearing attribute features, sample identity features, and sample extracted features, the training loss corresponding to the training sample images is calculated, and the training loss is used to train the wear detection model to obtain the trained wear detection model; by training the wear detection model from three aspects—wearing attributes, identity category, and image features—the training effect of the wear detection model is greatly improved, thereby improving the wear detection accuracy of the trained wear detection model.
[0037] Please see Figure 2 , Figure 2 This is a flowchart illustrating another embodiment of the training method for the wearable detection model provided in this application. The wearable detection model may include a feature extraction layer, an attribute prediction layer, and a category prediction layer. The method includes:
[0038] Step 21: Obtain training sample images.
[0039] Training sample images include target sample images, positive sample images, and negative sample images. Positive sample images can be images similar to target sample images, while negative sample images are images dissimilar to target sample images. The target sample image can be set according to the detection effect to be improved in this training. For example, if the current wearable detection model does not perform well in detecting fire suits when performing wearable detection tasks and easily confuses fire suits with other similar clothing, then images of fire suits can be selected as target sample images to improve the detection accuracy of the wearable detection model for fire suits through training.
[0040] A scheme for obtaining training sample images may include: obtaining a preset training sample set and target sample images, wherein the preset training sample set includes multiple sample image sets; and then selecting positive sample images and negative sample images from the multiple sample image sets based on the target sample images. Specifically, the preset training sample set may be a pre-set set of sample images used to train the wearable detection model, and each sample image set includes multiple sample images.
[0041] Furthermore, the wearing attribute label and identity label of each sample image in the preset training sample set can be obtained first. Then, sample images with the same identity label and the same wearing attribute label are grouped into the same sample image set, so that all sample images in the preset training sample set are divided into multiple sample image sets. The wearing attribute label and identity label of the sample image can be labeled by manual or model recognition methods, which is not limited here. In other embodiments, a sample wearing lookup table can be constructed based on the wearing attribute label and identity label of each sample image. The sample wearing lookup table contains each sample image and its corresponding wearing attribute label and identity label. Then, negative sample images and positive sample images that meet the conditions can be selected according to the sample wearing lookup table. This will not be described in detail here.
[0042] In one embodiment, the scheme for selecting positive and negative sample images from multiple sample image sets based on a target sample image may include: selecting a set of sample images from a preset training sample set that have the same wearing attribute features and identity features as the target sample image to obtain a positive sample image set; that is, the sample image set containing the target sample image can be regarded as the positive sample image set, and then a sample image is randomly selected from the positive sample image set to obtain a positive sample image. Selecting a set of sample images from the preset training sample set that have different identity features from the target sample image to obtain a negative sample image set; then, a sample image is randomly selected from the negative sample image set to obtain a negative sample image.
[0043] Understandably, at the start of each training session, a sample image can be randomly selected from the negative sample image set without replacement, and a sample image can be randomly selected from the positive sample image set without replacement. After the negative sample image set / positive sample image set has been traversed, the above method can be used to re-extract the negative sample image set / positive sample image from the preset training sample set, thereby ensuring that the images used in each training session are different and ensuring the model training effect.
[0044] In one embodiment, the negative sample image set may include a first negative sample image set and a second negative sample image set. When selecting negative sample images at the beginning of each training session, one sample image can be selected from each of the first and second negative sample image sets as a negative sample image. In this way, the amount of data in the negative sample image set can be increased, thereby improving the model training effect. The method for selecting the first and second negative sample image sets will be described in detail below.
[0045] A first negative sample image set can be obtained by selecting a first preset number of sample images from a preset training sample set that have the same wearing attribute features as the target sample image but different identity features. The wearing attribute can include multiple sub-attributes, the first preset number is the same as the number of sub-attributes, the wearing attribute features can include multiple sub-attribute features corresponding to the sub-attributes, and the wearing attribute label can include the sub-attribute label corresponding to each sub-attribute. Further, the multiple sub-attributes can include top style attributes, bottom style attributes, top color attributes, and bottom color attributes. The first negative sample image set can be obtained by selecting multiple sample image sets from the preset training sample set that have the same sub-attribute features as the target sample image but different identity features.
[0046] For example, if the identity label of the target sample image is "firefighter uniform", and the clothing attribute labels of the target sample image include the upper garment style attribute as short sleeve, the lower garment style attribute as long pants, the upper garment color attribute as red, and the lower garment color attribute as white, then four sample image sets can be selected from the preset training sample set as the first negative sample images. The four sample image sets are the sample image set with the upper garment style attribute as short sleeve and not a firefighter uniform, the sample image set with the lower garment style attribute as long pants and not a firefighter uniform, the sample image set with the upper garment color attribute as red and not a firefighter uniform, and the sample image set with the lower garment color attribute as white and not a firefighter uniform.
[0047] A second set of negative sample images can be obtained by randomly selecting a second preset number of sample images from the preset training sample set that have different identity features from the target sample image. In other words, at this time, it is not necessary to consider the wearing attribute label of the target sample image. The second set of negative sample images can be directly selected from the sample image set that has different identity features from the target sample image. For example, if the identity label of the target sample image is a fire suit, then a second preset number of sample images that are not fire suits can be randomly selected as the second negative sample image set. Understandably, the second preset number can be set according to the actual situation and is not limited here.
[0048] Step 22: Input the training sample images into the feature extraction layer to obtain the sample extracted features.
[0049] Step 22 is the same as step 12 in the above embodiments, and will not be repeated here.
[0050] Step 23: Input the sample extracted features into the attribute prediction layer to obtain the sample wearing attribute features.
[0051] Step 23 is the same as step 13 in the above embodiments, and will not be repeated here.
[0052] Step 24: Input the extracted features of the sample into the category prediction layer to obtain the sample identity features.
[0053] Step 24 is the same as step 14 in the above embodiments, and will not be repeated here.
[0054] After obtaining the sample clothing attribute features, sample identity features, and sample extraction features of the training sample images, the training loss corresponding to the training sample images can be calculated based on the sample clothing attribute features, sample identity features, and sample extraction features, as shown in steps 25 to 27 below.
[0055] Step 25: Calculate the attribute identity loss based on the sample's clothing attribute features and sample identity features.
[0056] The attribute-identity loss can include attribute loss values and identity loss values. First, the wearable attribute labels and identity labels corresponding to the training sample images are obtained; then, the loss between the sample's wearable attribute features and wearable attribute labels is calculated to obtain the attribute loss value; finally, the loss between the sample's identity features and identity labels is calculated to obtain the identity loss value. The wearable attribute labels and identity labels of the training sample images can be labeled manually or by model recognition, and are not limited here.
[0057] Furthermore, the sample wearing attribute features may include multiple sub-attribute features. Multiple cross-entropy loss values can be obtained by calculating the cross-entropy loss between each sub-attribute feature and its corresponding sub-attribute label. Then, the multiple cross-entropy loss values are weighted and summed to obtain the attribute loss value. The cross-entropy algorithm in this technical field can be used to calculate the cross-entropy loss between the sub-attribute features and their corresponding sub-attribute labels. No specific cross-entropy algorithm is limited here.
[0058] Similarly, the identity loss value can also be obtained by calculating the cross-entropy loss between the sample identity features and identity labels, which will not be elaborated here.
[0059] Step 26: Extract features based on the samples and calculate the feature loss value.
[0060] The features extracted from the samples can include the features of the target sample image, the features of the positive sample image, and the features of the negative sample image. Specifically, the triplet loss function can be used to calculate the features of the target sample image, the features of the positive sample image, and the features of the negative sample image to obtain the feature loss value. Calculating the feature loss value through the triplet loss function can increase the distance between the features of the positive sample image and the negative sample image, thereby improving the detection accuracy of the wearable detection model.
[0061] Specifically, the triplet loss function can be represented by the following formula (1):
[0062] L(A,P,N)=max(||f(A)-f(P)||-||f(A)-f(N)||+α,0) Formula (1)
[0063] In equation (1) above, L(A,P,N) represents the feature loss value, f(A) represents the feature of the target sample image, f(P) represents the feature of the positive sample image, f(N) represents the feature of the negative sample image, and α represents the hyperparameter.
[0064] Step 27: Calculate the training loss based on the attribute identity loss and feature loss values.
[0065] The training loss is obtained by weighted summation of the attribute loss value, identity loss value, and feature loss value; wherein, in a specific implementation, the weight ratio of the attribute loss value, identity loss value, and feature loss value can be 1:1:3.
[0066] Understandably, in other implementations, the weight ratios and specific weight values corresponding to the attribute loss value, identity loss value, and feature loss value can be set based on experience or actual conditions, and are not limited here. In addition, besides weighted summation of the attribute loss value, identity loss value, and feature loss value, other reasonable schemes can be used to calculate the training loss, such as multiplying the attribute loss value, identity loss value, and feature loss value to obtain the training loss.
[0067] In one embodiment, the output of the wearable detection model includes prediction results of attributes (denoted as attribute prediction results) and prediction results of clothing categories. The attribute prediction results include the color of the top and the color of the bottom, such as... Figure 3 As shown, the feature extraction layer includes a CNN network. The training sample images are input into the CNN network to obtain feature maps. The feature maps are reshaped to obtain sample extraction features. The sample extraction features are input into the attribute prediction layer and the category prediction layer respectively to obtain the output results. The output results include the color of the top being red, the color of the bottom being red, and the clothing category being petrochemical worker's clothing (i.e., the target object's identity is a petrochemical worker).
[0068] Step 28: Train the wear detection model based on the training loss to obtain the trained wear detection model.
[0069] Step 28 is the same as step 16 in the above embodiment, and will not be repeated here.
[0070] This embodiment selects positive and negative sample images from a preset training sample set based on the target sample image. The wearable detection model is trained using triplet images (including the target sample image, positive sample image, and negative sample image). This improves the detection accuracy of the wearable detection model for target sample images with poor detection performance, widens the distance between features of positive and negative sample images, effectively enhances the ability to distinguish negative sample images, and reduces confusion in the recognition of similar clothing. Furthermore, the training loss is obtained by weighted summation of triplet feature loss (i.e., feature loss value), attribute loss value, and identity loss value. This allows for training of the wearable detection model from three aspects: image features, attribute features, and identity category, ensuring the training effect of the wearable detection model and further improving the detection accuracy of the trained model.
[0071] Please see Figure 4 , Figure 4 This is a flowchart illustrating an embodiment of the wearable detection method provided in this application. The method includes:
[0072] Step 41: Obtain the image to be identified and the set of comparison images.
[0073] The image to be identified is the image for which the wear detection model is to be used for wear detection. The comparison image set may contain multiple comparison images and wear attribute labels and identity labels corresponding to each comparison image. Specifically, the comparison images can be used to perform feature comparison with the image to be identified, thereby determining whether the wear of the target object in the image to be identified meets the wear requirements corresponding to the identity label. For example, if the identity label of the comparison image is a fire suit, the wear detection method in this embodiment can be used to determine whether the target object in the image to be identified is wearing a fire suit in compliance with regulations.
[0074] Step 42: Input the image to be identified into the trained wearable detection model to obtain the features to be identified corresponding to the image.
[0075] The image to be identified can be input into the trained wearable detection model to obtain the features to be identified corresponding to the image. Specifically, the trained wearable detection model is obtained using the training method of the wearable detection model in the above embodiment. Feature extraction can be performed using the feature extraction layer in the wearable detection model to obtain the features to be identified corresponding to the image. For example, Figure 5 As shown, the feature extraction layer includes a CNN network. The image to be identified is input into the CNN network to obtain a feature map. The feature map is reshaped to obtain the features to be identified. The features to be identified are input into the attribute prediction layer to obtain the attribute prediction result.
[0076] Step 43: Input the comparison image into the trained wearable detection model to obtain the comparison features corresponding to the comparison image.
[0077] The comparison images are input into the trained wearable detection model, and the feature extraction layer in the wearable detection model is used to extract features, thereby obtaining the comparison features corresponding to the comparison images.
[0078] Step 44: Based on the features to be identified, the comparison features, and the wear attribute labels, obtain the wear detection results.
[0079] After extracting the features to be identified and the comparison features, the wear detection results can be obtained based on the features to be identified, the comparison features, and the wear attribute labels. Specifically, the wear detection results can include whether the wear of the target object in the image to be identified meets the wear requirements corresponding to the identity label.
[0080] In one specific embodiment, such as Figure 6 As shown, step 44 may include steps 61 to 64.
[0081] Step 61: Calculate the similarity between the feature to be identified and the corresponding feature of each image to obtain multiple similarities.
[0082] Calculate the similarity between the feature to be identified and the corresponding feature of each image to obtain multiple similarities; specifically, the cosine similarity between the feature to be identified and the image to be identified can be calculated to obtain the corresponding similarity. The specific calculation method of the similarity is not limited here.
[0083] Step 62: Calculate the maximum value of multiple similarities to obtain the maximum similarity.
[0084] The maximum similarity is obtained by selecting the maximum value from multiple similarity values; the image corresponding to the maximum similarity is the image that is most similar to the image to be identified.
[0085] Step 63: In response to the maximum similarity being greater than the first preset similarity threshold, the maximum similarity is adjusted based on the wearable attribute label to obtain the similarity adjustment value.
[0086] The system determines whether the maximum similarity score is greater than a first preset similarity threshold. If the maximum similarity score is less than or equal to the first preset similarity threshold, it indicates that the closest comparison image to the image to be identified also does not match the image to be identified, meaning the clothing worn by the target object does not meet the wearing requirements. If the maximum similarity score is greater than the first preset similarity threshold, the maximum similarity score can be adjusted based on the clothing attribute labels to obtain an adjusted similarity value. By adjusting the similarity score using clothing attributes and then using the adjusted similarity value to determine the clothing identification result, the recognition accuracy can be improved, and misjudgments caused by comparison images being similar clothing with the same color but different style, or the same style but different color, can be reduced. Understandably, the first preset similarity threshold can be set according to actual conditions and is not limited here.
[0087] Furthermore, the wearable attribute label includes the sub-attribute label corresponding to each sub-attribute and the confidence level. The confidence level product can be obtained by first calculating the product of the confidence levels corresponding to each sub-attribute label; then, the product of the confidence level product and the maximum similarity is calculated to obtain the similarity adjustment value. That is, the similarity adjustment value is calculated using the following formula (2):
[0088]
[0089] In equation (2) above, S1 represents the maximum similarity, and S2 represents the similarity adjustment value. Let N represent the confidence level of the i-th sub-attribute, and N represent the total number of sub-attributes.
[0090] Step 64: In response to the similarity adjustment value being greater than the second preset similarity threshold, determine that the wear detection result indicates that the wear of the target object meets the wear requirements.
[0091] If the similarity adjustment value is greater than the second preset similarity threshold, the wear detection result is determined to be that the wear of the target object meets the wear requirements; if the similarity adjustment value is less than or equal to the second preset similarity threshold, the wear detection result is determined to be that the wear of the target object does not meet the wear requirements. Understandably, the first preset similarity threshold can be set according to the actual situation, and is not limited here.
[0092] In this embodiment, when using the trained wear detection model for wear detection, the similarity can be adjusted according to the wear attribute labels, and then the adjusted similarity value is used to determine the wear recognition result. This can reduce the misjudgment caused by comparing images that are similar to the image to be identified but have the same color but different styles, or the same style but different colors, thereby improving the accuracy of wear detection.
[0093] Please see Figure 7 , Figure 7This is a schematic diagram of the structure of an embodiment of the model training device provided in this application. The model training device 70 includes a memory 71 and a processor 72 connected to each other. The memory 71 is used to store computer programs. When the computer programs are executed by the processor 72, they are used to implement the training method of the wearable detection model in the above embodiment.
[0094] Please see Figure 8 , Figure 8 This is a schematic diagram of an embodiment of the wearable detection device provided in this application. The wearable detection device 80 includes a memory 81 and a processor 82 connected to each other. The memory 81 is used to store computer programs. When the computer programs are executed by the processor 82, they are used to implement the wearable detection method in the above embodiment.
[0095] It is understood that the wearable detection device 80 may be the same device as the model training device in the above embodiments, and no limitation is made here.
[0096] Please see Figure 9 , Figure 9 This is a schematic diagram of an embodiment of a computer-readable storage medium provided in this application. The computer-readable storage medium 90 is used to store a computer program 91. When the computer program 91 is executed by a processor, it is used to implement the training method or wear detection method of the wear detection model in the above embodiment.
[0097] The computer-readable storage medium 90 can be any medium capable of storing program code, such as a server, USB flash drive, portable hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0098] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0099] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0100] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0101] If the technical solution of this application involves personal information, the product using this technical solution has clearly informed the user of the personal information processing rules and obtained the user's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using this technical solution has obtained the user's separate consent before processing the sensitive personal information, and also meets the requirement of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set up to inform users that they have entered the scope of personal information collection and that personal information will be collected. If an individual voluntarily enters the collection scope, it is deemed that they have agreed to the collection of their personal information; or on the personal information processing device, with clear signs / information informing users of the personal information processing rules, authorization is obtained from the individual through pop-up information or by asking the individual to upload their personal information; wherein, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.
[0102] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A training method for a wearable detection model, characterized in that, The wearable detection model includes a feature extraction layer, an attribute prediction layer, and a category prediction layer; the method includes: Obtain training sample images; The training sample images are input into the feature extraction layer to obtain the sample extracted features; The extracted features from the sample are input into the attribute prediction layer to obtain the sample's wear attribute features; The extracted features from the sample are input into the category prediction layer to obtain the sample identity features; Based on the sample wear attribute features, the sample identity features, and the sample extraction features, calculate the training loss corresponding to the training sample image; Based on the training loss, the wear detection model is trained to obtain the trained wear detection model; The training sample images include target sample images, positive sample images, and negative sample images. The step of obtaining the training sample images includes: Obtain a preset training sample set and the target sample image, wherein the preset training sample set includes multiple sample image sets; Based on the target sample image, the positive sample image and the negative sample image are selected from the plurality of sample image sets; The step of selecting the positive sample image and the negative sample image from the plurality of sample image sets based on the target sample image includes: A set of positive sample images is obtained by selecting a set of sample images from the preset training sample set that have the same wearing attribute features and identity features as the target sample image; A sample image is randomly selected from the set of positive sample images to obtain the positive sample image; A negative sample image set is obtained by selecting a set of sample images from the preset training sample set that have different identity features from the target sample image; The negative sample image is obtained by randomly selecting one sample image from the set of negative sample images. The negative sample image set further includes a first negative sample image set and a second negative sample image set; the step of selecting a set of sample images with different identity features from the target sample image from the preset training sample set to obtain the negative sample image set includes: The first negative sample image set is obtained by selecting a first preset number of sample images from the preset training sample set that have the same wearing attribute features as the target sample image but different identity features; The second negative sample image set is obtained by randomly selecting a second preset number of sample images from the preset training sample set that have different identity features from the target sample image; The wearable attribute features include multiple sub-attribute features; the step of selecting a first preset number of sample images from the preset training sample set that have the same wearable attribute features as the target sample image but different identity features to obtain the first negative sample image set includes: Multiple sample image sets that are identical to each sub-attribute feature of the target sample image but different in identity feature are selected from the preset training sample set to obtain the first negative sample image set.
2. The training method for the wearable detection model according to claim 1, characterized in that, The step of calculating the training loss corresponding to the training sample image based on the sample wear attribute features, the sample identity features, and the sample extraction features includes: Based on the sample's wear attribute features and the sample's identity features, calculate the attribute identity loss; Based on the extracted features from the samples, the feature loss value is calculated; The training loss is calculated based on the attribute identity loss and the feature loss value.
3. The training method for the wearable detection model according to claim 2, characterized in that, The attribute identity loss includes an attribute loss value and an identity loss value. The step of calculating the attribute identity loss based on the sample's wearable attribute features and the sample's identity features includes: Obtain the wearable attribute labels and identity labels corresponding to the training sample images; Calculate the loss between the sample wearing attribute features and the wearing attribute labels to obtain the attribute loss value; Calculate the loss between the sample identity features and the identity label to obtain the identity loss value; The step of calculating the training loss based on the attribute identity loss and the feature loss value includes: The training loss is obtained by weighted summation of the attribute loss value, the identity loss value, and the feature loss value.
4. The training method for the wearable detection model according to claim 3, characterized in that, The sample wearable attribute features include multiple sub-attribute features, and the wearable attribute labels include sub-attribute labels corresponding to each sub-attribute. The step of calculating the loss between the sample wearable attribute features and the wearable attribute labels to obtain the attribute loss value includes: Calculate the cross-entropy loss between each sub-attribute feature and its corresponding sub-attribute label to obtain multiple cross-entropy loss values; The attribute loss value is obtained by weighted summation of the multiple cross-entropy loss values.
5. The training method for the wearable detection model according to claim 2, characterized in that, The training sample images include target sample images, positive sample images, and negative sample images; the extracted sample features include features of the target sample images, features of the positive sample images, and features of the negative sample images; the step of calculating the feature loss value based on the extracted sample features includes: The feature loss value is obtained by calculating the features of the target sample image, the positive sample image, and the negative sample image using a triplet loss function.
6. A wearable detection method, characterized in that, include: Obtain the image to be identified and the set of comparison images; The comparison image set includes multiple comparison images and wearable attribute tags and identity tags corresponding to each comparison image; The image to be identified is input into the trained wearable detection model to obtain the features to be identified corresponding to the image to be identified. The trained wearable detection model is trained using the training method of the wearable detection model in any one of claims 1-5. The comparison image is input into the trained wearable detection model to obtain the comparison features corresponding to the comparison image; Based on the features to be identified, the comparison features, and the wear attribute labels, a wear detection result is obtained; the wear detection result includes whether the wear of the target object in the image to be identified meets the wear requirements corresponding to the identity label; The step of obtaining the wear detection result based on the feature to be identified, the comparison feature, and the wear attribute label includes: Calculate the similarity between the feature to be identified and the corresponding feature of each comparison image to obtain multiple similarities; Calculate the maximum value of the multiple similarities to obtain the maximum similarity; In response to the maximum similarity being greater than a first preset similarity threshold, the maximum similarity is adjusted based on the wearable attribute tag to obtain a similarity adjustment value; In response to the similarity adjustment value being greater than a second preset similarity threshold, the wear detection result is determined to be that the wear of the target object meets the wear requirements; The wearable attribute label includes a sub-attribute label and a confidence score for each sub-attribute; the step of adjusting the maximum similarity based on the wearable attribute label to obtain an adjusted similarity value includes: Calculate the product of the confidence scores corresponding to each of the sub-attribute labels to obtain the confidence score product; The product of the confidence score and the maximum similarity is calculated to obtain the similarity adjustment value.
7. A model training device, characterized in that, It includes an interconnected memory and a processor, wherein the memory is used to store a computer program, which, when executed by the processor, is used to implement the training method of the wearable detection model according to any one of claims 1-5.
8. A wearable detection device, characterized in that, It includes an interconnected memory and a processor, wherein the memory is used to store a computer program, which, when executed by the processor, is used to implement the wearable detection method of claim 6.
9. A computer-readable storage medium for storing a computer program, characterized in that, When executed by a processor, the computer program is used to implement the training method of the wearable detection model according to any one of claims 1-5 or the wearable detection method according to claim 6.
Citation Information
Patent Citations
Detection method and detection device
CN107609544A
Object recognition model training method, object recognition method and corresponding devices
CN110414432A
Target recognition model training method and device and electronic equipment
CN112990432A