Cleaning and care mode determination method based on multi-modal fusion and electronic device
By adopting a multimodal fusion method in the water-washed mark recognition, using the target fusion model to fusion images and text features, the problem of poor water-washed mark recognition effect in the single-mode architecture is solved, and efficient identification of water-washed marks is achieved.
Patent Information
- Application Number
- CN202311687380.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-08
- Publication Date
- 2025-06-10
AI Technical Summary
In the prior art, since the identification of the water-washed mark adopts a single-mode architecture and there is a lot of unrelated information on the water-washed mark, the identification effect of the water-washed mark is poor.
The washing and care method based on multimodal fusion is adopted to feature the text features and image features in the image to be identified through the target fusion model to obtain the recognition results. The method includes icon and text annotation on the image, generating an image data set using a preset annotation method, and extracting text and image features through the training target fusion model for weighted fusion.
The recognition effect of washing marks has been improved, and the problem of poor recognition effect of washing marks in single-mode architecture is overcome, especially in small samples and multi-sample data sets.
Smart Images

Figure CN120125866A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of smart home. Specifically, it relates to a method for determining a washing and care method based on multi-modal fusion and an electronic device. Background Art
[0002] The development of technologies such as computer vision-based clothing recognition, attribute recognition, and virtual fitting has promoted the development of clothing digitization and made great contributions to the clothing industry. However, there is a gap in the aspect of clothing washing and care methods.
[0003] In related technologies, one recognition method is to use clothing material recognition based on Radio Frequency Identification (RFID) tags to match washing and care methods through materials. However, this method requires the clothing itself to carry RFID tags and the cooperation of the upstream and downstream industrial chains.
[0004] There is also a recognition method that uses clothing wash labels for clothing recognition. However, most of the existing clothing wash label recognitions are trained using supervised learning methods and recognized within a limited training set to achieve ideal results within a certain closed set range. However, the problems of this detection-based method are as follows: 1) The washing and care icon standards of wash labels in different countries are different, and the supervised training data has a long-tail distribution. Icons of small samples and zero samples cannot be recognized; 2) After multiple washes, the icons on the wash labels themselves have problems such as fading, wrinkling, and damage, affecting the recognition effect; 3) As the number of label categories increases, the performance of the detector also decreases; 4) There is a large amount of irrelevant text information on the wash label; 5) The classification-based method classifies the entire image or region and is not restricted by the number of categories, but it relies heavily on region proposals. When there is a complex background in the image, the algorithm performance will be affected.
[0005] In view of the problems in related technologies, such as the use of a single-mode architecture for the recognition of wash labels and the existence of a large amount of irrelevant information on the wash labels, resulting in poor recognition effects of wash labels, no effective solutions have been proposed yet.
[0006] Therefore, it is necessary to improve related technologies to overcome the above-mentioned defects in related technologies. Summary of the Invention
[0007] Embodiments of this application provide a method for determining a washing and care method based on multi-modal fusion and an electronic device, so as to at least solve the problem in related technologies that due to the use of a single-mode architecture for the recognition of wash labels and the existence of a large amount of irrelevant information on the wash labels, the recognition effect of wash labels is poor.
[0008] According to one aspect of the embodiments of the present application, a method for determining a washing and care method based on multi-modal fusion is provided, including: inputting a first image to be recognized into a target fusion model; wherein, the target fusion model is trained by a first data set, and the first data set includes: image data obtained by performing icon and text annotation on a second image according to a preset annotation method; extracting text features corresponding to the first image through a first model branch of the target fusion model, and extracting image features corresponding to the first image through a second model branch of the target fusion model; performing a feature fusion operation on the text features and the image features according to a preset fusion method to obtain a recognition result of the first image, and determining the recognition result as the washing and care method of the target clothing, wherein the target clothing carries a washing label corresponding to the first image.
[0009] In an exemplary embodiment, before inputting the first image to be recognized into the target fusion model, the method further includes: performing a hit operation on the icons and text in the second image to obtain the hit icons and the hit text: determining a first label corresponding to the hit icons in a preset multi-level label system, and annotating the hit icons with the description text indicated by the first label to obtain the icon annotation data of the second image; and recognizing first text information of the hit text according to a preset recognition method, and annotating the hit text with the first text information to obtain the text annotation data of the second image; wherein, the first label is the label with the highest similarity to the hit icons in the multi-level label system; determining the second image, the icon annotation data and the text annotation data as the content included in the first data set.
[0010] In an exemplary embodiment, before inputting the first image to be recognized into the target fusion model, the method further includes: performing a hit operation on the icons and text in the third image to obtain the icons of the third image after hitting and the text of the third image after hitting, where the washing label corresponding to the third image is different from the washing label corresponding to the second image; determining the second label corresponding to the icons of the third image after hitting in a preset multi-level label system, and annotating the icons of the third image after hitting with the description text indicated by the second label to obtain the icon annotation data of the third image; and identifying the second text information of the text of the third image after hitting according to a preset recognition method, and annotating the text of the third image after hitting with the second text information to obtain the text annotation data of the third image; where the second label is the label with the highest similarity to the text of the third image after hitting in the multi-level label system; determining the third image, the icon annotation data of the third image, and the text annotation data of the third image as the content included in the second data set; where the second data set is used to determine the model architecture of the target fusion model.
[0011] In an exemplary embodiment, after annotating the icons in the hit third image and the text in the third image in the following manner to obtain the second data set, the method further includes: determining the label detection network allowed to be adopted by the second model branch, and determining multiple feature extraction networks allowed to be used for feature extraction of the labels detected by the label detection network; inputting the second data set into multiple simulation branches respectively to obtain the first precision of the multiple simulation branches for feature extraction of the second data set respectively, where the multiple simulation branches are simulation branches constructed by the label detection network and the multiple feature extraction networks respectively; and determining the simulation branch corresponding to the highest first precision among the multiple first precisions as the second model branch.
[0012] In an exemplary embodiment, after annotating the icons in the hit third image and the text in the third image in the following manner to obtain the second data set, the method further includes: when the second model branch has been determined and M preset networks have been preset for the first model branch, determining the preset recognition order of the M preset networks for the first graphic in the first model branch, where M is a positive integer; performing ablation experiments on the M preset networks in reverse order according to the preset recognition order to obtain M experimental models corresponding to the target fusion model; inputting the second data set into the M experimental models respectively to determine the second precision of the M experimental models for image recognition of the second data set respectively; and determining the experimental model corresponding to the highest second precision among the multiple second precisions as the target fusion model.
[0013] In an exemplary embodiment, after labeling the icons in the hit third image and the text in the third image to obtain a second data set, the method further includes: determining multiple feature fusion methods preset for the multi-modal fusion space included in the target fusion model, where the multi-modal fusion space is used to perform feature fusion on the text features extracted by the first model branch included in the target fusion model and the image features extracted by the second model branch; inputting the second data set into the target fusion model adopting different feature fusion methods respectively to determine the third accuracy of the target fusion model adopting different feature fusion methods for image recognition, where the different feature fusion methods are different feature fusion methods included in the multiple feature fusion methods; determining the feature fusion method corresponding to the highest third accuracy among the multiple third accuracies as the preset fusion method.
[0014] In an exemplary embodiment, extracting the text features corresponding to the first image through the first model branch of the target fusion model includes: in the case where it is determined that the first model branch includes one or a combination of a character recognition network, a text encoding network, and a recurrent neural network, converting the text image in the first image into editable text through the character recognition network; encoding the editable text into a corresponding text sequence through the text encoding network; modeling the text sequence through the recurrent neural network to extract the text features in the text sequence.
[0015] In an exemplary embodiment, performing a feature fusion operation on the text features and the image features according to the preset fusion method to obtain an identification result of the first image includes: in the case where the preset fusion method is weighted fusion, determining a first clarity degree of the icons in the first image and a second clarity degree of the text in the first image; determining a first fusion weight corresponding to the image features and a second fusion weight corresponding to the text features according to the proportional relationship between the first clarity degree and the second clarity degree; determining the identification result through the image features, the text features, the first fusion weight, and the second fusion weight.
[0016] According to another aspect of the embodiments of the present application, there is also provided a device for determining a washing and care method based on multimodal fusion, including: an input module, configured to input a first image to be recognized into a target fusion model; wherein, the target fusion model is trained through a first data set, and the first data set includes: image data obtained by performing icon and text annotation on a second image according to a preset annotation method; an extraction module, configured to extract text features corresponding to the first image through a first model branch of the target fusion model, and extract image features corresponding to the first image through a second model branch of the target fusion model; a fusion module, configured to perform a feature fusion operation on the text features and the image features according to a preset fusion method to obtain a recognition result of the first image, and determine the recognition result as the washing and care method of a target garment, wherein the target garment carries a washing label corresponding to the first image.
[0017] According to still another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the above-mentioned method for determining a washing and care method based on multimodal fusion when running.
[0018] According to still another aspect of the embodiments of the present application, there is also provided an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the above-mentioned processor executes the above-mentioned method for determining a washing and care method based on multimodal fusion through the computer program.
[0019] Through the present application, a first image to be recognized is input into a target fusion model; wherein, the target fusion model is trained through a first data set, and the first data set includes: image data obtained by performing icon and text annotation on a second image according to a preset annotation method; text features corresponding to the first image are extracted through a first model branch of the target fusion model, and image features corresponding to the first image are extracted through a second model branch of the target fusion model; a feature fusion operation is performed on the text features and the image features according to a preset fusion method to obtain a recognition result of the first image, and the recognition result is determined as the washing and care method of a target garment, wherein the target garment carries a washing label corresponding to the first image. By adopting the above technical solution, the problem in the related art that the recognition effect of the washing label is poor due to the use of a single-mode architecture for recognizing the washing label and a large amount of irrelevant information exists on the washing label is solved; thereby improving the recognition effect of the washing mark. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0021] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0022] Figure 1 It is a schematic diagram of the hardware environment of an optional method for determining a washing and care method based on multimodal fusion according to an embodiment of the present application;
[0023] Figure 2 It is a flowchart of an optional method for determining a washing and care method based on multimodal fusion according to an embodiment of the present application;
[0024] Figure 3 It is a schematic diagram (one) of the washing label annotation of an optional method for determining a washing and care method based on multimodal fusion according to an embodiment of the present application;
[0025] Figure 4 It is a schematic diagram (two) of the washing label annotation of an optional method for determining a washing and care method based on multimodal fusion according to an embodiment of the present application;
[0026] Figure 5 It is a schematic diagram of the label system of an optional method for determining a washing and care method based on multimodal fusion according to an embodiment of the present application;
[0027] Figure 6 It is a schematic diagram (one) of the icon of an optional method for determining a washing and care method based on multimodal fusion according to an embodiment of the present application;
[0028] Figure 7 It is a schematic diagram (two) of the icon of an optional method for determining a washing and care method based on multimodal fusion according to an embodiment of the present application;
[0029] Figure 8 It is a model framework diagram of an optional method for determining a washing and care method based on multimodal fusion according to an embodiment of the present application;
[0030] Figure 9 It is a structural block diagram of an optional device for determining a washing and care method based on multimodal fusion according to an embodiment of the present application. Detailed implementation manners
[0031] To enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solution in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.
[0032] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0033] According to one aspect of the embodiments of this application, a method for determining a washing and care method based on multimodal fusion is provided. The method for determining a washing and care method based on multimodal fusion is widely applied to whole-house intelligent digital control application scenarios such as Smart Home, smart home, smart home appliance ecosystem, and IntelligenceHouse ecosystem. Optionally, in this embodiment, the above-mentioned method for determining a washing and care method based on multimodal fusion can be applied to, for example Figure 1 the hardware environment composed of multiple terminal devices 102 and a server 104 as shown. As Figure 1 shown, the server 104 is connected to multiple terminal devices 102 through a network, and can be used to provide services (such as application services, etc.) for the terminal or the client installed on the terminal. A database can be set on the server or independently of the server, and is used to provide data storage services for the server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server, and are used to provide data operation services for the server 104.
[0034] The above network may include, but is not limited to, at least one of the following: a wired network, a wireless network. The above wired network may include, but is not limited to, at least one of the following: a wide area network, a metropolitan area network, a local area network. The above wireless network may include, but is not limited to, at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The terminal device 102 may not be limited to a PC, a mobile phone, a tablet computer, a smart air conditioner, a smart range hood, a smart refrigerator, a smart oven, a smart stove, a smart washing machine, a smart water heater, a smart washing device, a smart dishwasher, a smart projection device, a smart TV, a smart drying rack, a smart curtain, a smart audio and video system, a smart socket, a smart speaker, a smart sound box, a smart fresh air device, a smart kitchen and bathroom device, a smart bathroom device, a smart floor sweeping robot, a smart window cleaning robot, a smart mopping robot, a smart air purification device, a smart steam box, a smart microwave oven, a smart kitchen water heater, a smart purifier, a smart water dispenser, a smart door lock, etc.
[0035] In this embodiment, a method for determining a washing and care method based on multi-modal fusion is provided, including but not limited to being applied to a smart speaker. Figure 2 It is a flowchart of the method for determining a washing and care method based on multi-modal fusion according to an embodiment of the present application. The process includes the following steps:
[0036] Step S202: Input the first image to be recognized into the target fusion model; wherein, the target fusion model is trained by a first data set, and the first data set includes: image data obtained by performing icon and text annotation on a second image according to a preset annotation method.
[0037] Step S204: Extract the text features corresponding to the first image through the first model branch of the target fusion model, and extract the image features corresponding to the first image through the second model branch of the target fusion model.
[0038] Step S206: Perform a feature fusion operation on the text features and the image features according to a preset fusion method to obtain the recognition result of the first image, and determine the recognition result as the washing and care method of the target clothing, where the target clothing carries a washing label corresponding to the first image.
[0039] Through the above steps, the first image to be recognized is input into the target fusion model; wherein, the target fusion model is trained by a first data set, and the first data set includes: image data obtained by labeling icons and texts on a second image according to a preset labeling method; extracting text features corresponding to the first image through a first model branch of the target fusion model, and extracting image features corresponding to the first image through a second model branch of the target fusion model; performing a feature fusion operation on the text features and the image features according to a preset fusion method to obtain a recognition result of the first image, and determining the recognition result as the washing method of the target clothing, wherein the target clothing carries a washing label corresponding to the first image. By adopting the above technical solution, the problem in the related art that the recognition effect of the washing label is poor due to the use of a single-mode architecture for the recognition of the washing label and the existence of a large amount of irrelevant information on the washing label is solved; thus, the recognition effect of the washing mark is improved.
[0040] Before performing the above step S202, it is necessary to determine the target fusion data. Therefore, before inputting the first image to be recognized into the target fusion model, the method further includes: performing a hit operation on the icons and texts in the second image to obtain the hit icons and the hit texts: determining a first label corresponding to the hit icon in a preset multi-level label system, and labeling the hit icon with the description text indicated by the first label to obtain the icon labeling data of the second image; and identifying the first text information of the hit text according to a preset recognition method, and labeling the hit text with the first text information to obtain the text labeling data of the second image; wherein, the first label is the label with the highest similarity to the hit icon in the multi-level label system; determining the second image, the icon labeling data and the text labeling data as the content included in the first data set.
[0041] It can be understood that the first image is the graph of the washing mark (i.e., the washing label) to be recognized. The hit operation can be understood as selecting the washing method label (i.e., the hit icon) and the washing annotation text (i.e., the hit text) in the first image through a closed rectangular box. For example, "dry cleaning", "hanging to dry", etc., as specifically shown in Figure 3 and Figure 4 shown.
[0042] In the embodiments of the present application, referring to the standards of different countries, the labels included in the washing mark are classified to construct a multi-level label system. The multi-level label system includes a first-level label and a second-level label. The first-level labels are: washing, bleaching, drying, ironing, professional textile maintenance. The second-level labels are further subdivided on the basis of the first-level labels. For example, for the washing category, there are different washing methods, etc., as specifically shown in Figure 5 shown.
[0043] During the annotation process, the corresponding description text is annotated beside the hit icon. For example, Figure 3 beside the third positive icon hit by the rectangular frame in [Figure], there is a label of "Flat drying in the sun". The corresponding text information is annotated beside the hit text. For example, Figure 3 beside "Flat drying" hit by the rectangular frame in [Figure], there is a label of "Flat drying". It should be noted that all icons in the care label image and text information related to washing and care need to be hit and annotated.
[0044] Optionally, before inputting the first image to be recognized into the target fusion model, it is also necessary to determine the model architecture of the target fusion model. Specifically, a second dataset can be obtained by annotating the care label image for determining the model architecture. The method further includes: performing a hit operation on the icons and text in the third image to obtain the icons of the hit third image and the text of the hit third image, where the care label corresponding to the third image is different from the care label corresponding to the second image; determining the second label corresponding to the icons of the hit third image in a pre-set multi-level label system, and annotating the icons of the hit third image with the description text indicated by the second label to obtain the icon annotation data of the third image; and recognizing the second text information of the text of the hit third image according to a pre-set recognition method, and annotating the text of the hit third image with the second text information to obtain the text annotation data of the third image; where the second label is the label with the highest similarity to the text of the hit third image in the multi-level label system; determining the third image, the icon annotation data of the third image, and the text annotation data of the third image as the content included in the second dataset; where the second dataset is used to determine the model architecture of the target fusion model.
[0045] The annotation process of the second dataset is similar to that of the first dataset, and will not be elaborated in this embodiment of the present application. It should be noted that the first dataset can also be used in the process of determining the model architecture.
[0046] After determining the second data set, further, after annotating the icons in the hit third image and the text in the third image in the following manner to obtain the second data set, the method further includes: determining the label detection network allowed to be adopted by the second model branch, and determining a plurality of feature extraction networks allowed to be used for feature extraction of the labels detected by the label detection network; inputting the second data set into a plurality of simulation branches respectively to obtain the first precision of the plurality of simulation branches for feature extraction of the second data set respectively, where the plurality of simulation branches are simulation branches constructed by respectively combining the label detection network with the plurality of feature extraction networks; and determining the simulation branch corresponding to the highest first precision among the plurality of first precisions as the second model branch.
[0047] The second model branch can be understood as a branch for feature extraction of the icons in the wash mark. The label detection network uses Yolov5, and detects the extraction precision of a plurality of feature extraction networks that can be used in conjunction with Yolov5, and constructs the second model branch by combining the feature extraction network with the highest precision with Yolov5. The feature extraction networks that can be used in conjunction with Yolov5 include, but are not limited to: MobileNetv3, InceptionV3, etc.
[0048] After determining the second data set by annotating the icons in the hit third image and the text in the third image in the following manner, the method further includes: when the second model branch has been determined and M preset networks have been preset for the first model branch, determining the preset recognition order of the M preset networks for the first graphic in the first model branch, where M is a positive integer; performing ablation experiments on the M preset networks in reverse order according to the preset recognition order to obtain M experimental models corresponding to the target fusion model; inputting the second data set into the M experimental models respectively to determine the second precision of the M experimental models for image recognition of the second data set respectively; and determining the experimental model corresponding to the highest second precision among the plurality of second precisions as the target fusion model.
[0049] The M preset networks can be, in sequence: OCR, IF-IDF, BERT+LSTM, that is, M = 3 at this time. Ablation experiments are performed on these three preset networks in reverse order. Among them, the first branches of the 3 experimental models are: OCR+IF-IDF+BERT+LSTM, OCR+IF-IDF, OCR. Therefore, the experimental model with the highest image recognition precision is selected as the final model architecture.
[0050] It should be noted that OCR (Optical Character Recognition) is an optical character recognition algorithm used to convert text in images into editable text.
[0051] IF-IDF (Inverse Document Frequency-Inverse Term Frequency) is an algorithm used for information retrieval and text mining to evaluate the importance of a word for a document set or corpus.
[0052] BERT (Bidirectional Encoder Representations from Transformers) is a natural language processing algorithm based on the Transformer model used for text representation learning and language understanding tasks.
[0053] LSTM (Long Short-Term Memory) is a variant of the Recurrent Neural Network (RNN) used to process and predict time series data.
[0054] Further, after annotating the icons and text in the hit third image in the following manner to obtain the second data set, the method further includes: determining multiple feature fusion methods preset for the multi-modal fusion space included in the target fusion model, where the multi-modal fusion space is used to perform feature fusion on the text features extracted by the first model branch and the image features extracted by the second model branch included in the target fusion model; inputting the second data set into the target fusion models using different feature fusion methods respectively to determine the third accuracy of the target fusion models using different feature fusion methods for image recognition, where the different feature fusion methods are different feature fusion methods included in the multiple feature fusion methods; and determining the feature fusion method corresponding to the highest third accuracy among the multiple third accuracies as the preset fusion method.
[0055] The fusion methods for performing feature fusion on image features and text features include: summation fusion, average fusion, weighted fusion, and concatenation fusion. In the embodiments of the present application, according to the contribution accuracy (i.e., the third accuracy) of these fusion methods to the final image recognition, the fusion method with the highest third accuracy is used as the preferred fusion method.
[0056] However, in an alternative embodiment, multiple parallel sub-fusion spaces may be set up in the multi-modal fusion space to perform feature fusion on the obtained image features and text features respectively and simultaneously, and summarize the multiple sub-recognition results finally output. The recognition result that appears most frequently among the multiple sub-recognition results is used as the final recognition result to improve the accuracy of the washing label recognition.
[0057] After determining the target fusion model architecture, in an alternative embodiment, extracting the text features corresponding to the first image through the first model branch of the target fusion model includes: in the case where it has been determined that the first model branch includes one or several combinations of a character recognition network, a text encoding network, and a recurrent neural network, converting the text image in the first image into editable text through the character recognition network; encoding the editable text into a corresponding text sequence through the text encoding network; and modeling the text sequence through the recurrent neural network to extract the text features in the text sequence.
[0058] In an exemplary embodiment, performing a feature fusion operation on the text features and the image features according to a preset fusion method to obtain a recognition result for the first image includes: in the case where the preset fusion method is weighted fusion, determining a first clarity degree of the icon in the first image and determining a second clarity degree of the text in the first image; determining a first fusion weight corresponding to the image features and a second fusion weight corresponding to the text features according to the proportional relationship between the first clarity degree and the second clarity degree; and determining the recognition result through the image features, the text features, the first fusion weight, and the second fusion weight.
[0059] Optionally, determining the recognition result through the image features, the text features, the first fusion weight, and the second fusion weight includes: determining a first product of the first fusion weight and the image features and determining a second product of the second fusion weight and the text features; and determining the sum of the first product and the second product as the recognition result.
[0060] The first image can be divided into an icon part and a text part, so as to determine the clarity degrees of the icon and the text.
[0061] Specifically, the clarity can be evaluated in the following ways: 1) Clarity evaluation algorithm based on local image features: This algorithm evaluates the clarity of an image by analyzing local features of the image, such as edges, textures, etc. Common methods include using gradient information, contrast information, etc. to calculate clarity metrics. 2) Clarity evaluation algorithm based on frequency domain analysis: This algorithm evaluates the clarity of an image by performing frequency domain analysis on the image, such as Fourier transform. A clear image usually has more high-frequency components. 3) Clarity evaluation algorithm based on deep learning: Train a deep neural network to learn the clarity features of the image.
[0062] The method for determining the fusion weights is not limited to clarity and can also be determined by feedback. For example, the initial weights are set to 5:5. In the case of receiving an incorrect recognition result, extract the text features and image features corresponding to the incorrect result. When it is determined that either the text features or the image features corresponding to the incorrect result have a greater correlation with the correct result, increase the weight ratio corresponding to that feature. The specific degree of increase can be set as needed. For example, the original weight ratio is text feature: image feature = 5:5; when it is determined that the text feature is the feature, change the weight ratio to 6:4.
[0063] Without determining the weights according to clarity, the image can also be preprocessed, such as denoising, sharpening, etc. to improve the image clarity.
[0064] Obviously, the above-described embodiments are only a part of the embodiments of the present application, rather than all embodiments. To better understand the above method for determining the washing and care method based on multi-modal fusion, the following will illustrate the above process in combination with embodiments, but it is not used to limit the technical solutions of the embodiments of the present application. Specifically:
[0065] The objectives of the optional embodiments of the present application are as follows:
[0066] There are image features and a large amount of text information in the wash label. For the wash label recognition task, it is found that the visual space and the text space promote each other during the category retrieval process. For the small sample problem, additional text information can improve the metrics. Therefore, the optional embodiments of the present application propose an end-to-end, multi-task learning architecture based on multi-modal fusion, using the method of multi-modal fusion of the image space and the text space, and adopting a scheme combining image detection and classification to recognize the wash label. Experiments prove that the proposed scheme has a leading advantage compared with single-modal recognition, with a 7.29% improvement in performance, and also shows obvious advantages on small sample datasets.
[0067] Specifically, the wash label recognition scheme of the optional embodiments of the present application is as follows:
[0068] An alternative embodiment of the present application proposes a label hierarchy division system for clothing washing and care methods (equivalent to the multi-level label system in the above embodiment), and proposes a method for multi-modal data annotation of washing labels with text annotations and label markings based on vision. The data set is divided into first-level labels and second-level labels according to the secondary classification method. As Figure 5 shown, the first-level labels are: washing, bleaching, drying, ironing, professional textile maintenance, and others. When dividing, not only the standards of multiple countries are referred to, but six first-level categories are also divided according to image features. The washing tank format represents the washing category; the triangle represents the bleaching category; the square represents the drying category; the manual iron represents the ironing category; the circle represents the maintenance category of professional dry cleaning and professional wet cleaning; those that do not conform to the above category features are classified as other categories. As Figure 5 shown, the second-level labels are further divided on the basis of the first-level labels. For example, for the first-level label of the washing category, there are different washing methods, including but not limited to: non-hand wash; for the first-level label of the bleaching category, there are different bleaching methods, including but not limited to: no bleaching; for the first-level label of the drying category, there are different drying methods, including but not limited to: drying on a line, drying flat; for the first-level label of the ironing category, there are different ironing methods, including but not limited to: ironing flat at 110 degrees; for the first-level label of professional textile maintenance, there are different maintenance methods, including but not limited to: no professional dry cleaning required; for the first-level label of other categories, there are different second-level labels, including but not limited to: neutral washing.
[0069] Based on the above label hierarchy division system, the embodiment of the present application also provides a graphic and text annotation method for washing labels, including: 1) Icon annotation method: The first-level label categories are annotated from the picture according to the detected bounding box. The second-level sub-categories are classified by small icons obtained by cropping the first-level category labels according to the bounding box. 2) Text annotation method: The part of the washing label related to the washing and care method is selected, and the position of the corresponding text description (equivalent to the hit text in the above embodiment) and the content of the text (equivalent to the text annotation data in the above embodiment) are annotated. Specifically, as Figure 3 and Figure 4 shown
[0070] 3) As for the problem of the diversity of wash label images, the image features of wash labels with different image features of the same category are aligned with the text features through weak label marking. This problem is solved by building a one-to-many mapping relationship between text and image. We marked the correspondence between icons and text. There are multiple icons corresponding to the same text category description. For example, taking the conventional wash at 30℃ as an example, the image features will be as follows Figure 6 The difference is shown in the figure, but the corresponding text description is "the maximum washing temperature is 30 degrees". Taking the ironing 150℃ label as an example, different image features such as Figure 7 As shown, but corresponding to the same article description "maximum ironing is 150 degrees".
[0071] Furthermore, an optional embodiment of the present application also provides a method for image-text fusion recognition of wash labels. The method provides a wash label recognition framework (equivalent to the target fusion model in the above embodiment). The framework is end-to-end and does not require additional input. Only the wash label image needs to be input to obtain the label recognition result. Figure 8 As shown, the final framework consists of two branches, a visual feature extraction branch (equivalent to the second model branch in the above embodiment) and a text feature extraction branch (equivalent to the first model in the above embodiment). Finally, feature encoding fusion is performed in the image-text shared representation space (equivalent to the multi-modal fusion space in the above embodiment). Finally, the vectors of the unified modality are classified to obtain the recognition result.
[0072] The optional embodiment of the present application trains all models on an NVIDIA Tesla K80 GPU with 11GB of video memory. First, we selected Yolov5 as the region extraction target detection network. On this basis, we tested the performance of the currently commonly used feature extraction networks on the data set. The specific indicators are shown in Table 1. Finally, Yolov5+InceptionV3 was selected as the baseline of the image branch.
[0073] Table 1 Performance of image branch on datasets
[0074] Only Image (Yolov5) Accuracy +MobileNetv3 79.94 +InceptionV3 80.43 +ResNet50 70.97 +VGG16 65.49 +VGG19 68.87 +Xception 78.74 +DenseNet121 74.95 +NASNet 79.19 +MobileNetV2 70.82 +InceptionResNetV2 77.56
[0075] After the baseline of the image branch is determined, an ablation experiment is performed on the text branch. The effects of OCR, IF-IDF, and BERT+LSTM algorithms on recognition performance are verified respectively. The specific performance is shown in Table 2. The final accuracy of multi-modal fusion is 7.29% higher than that of single-mode image recognition.
[0076] Table 2 Performance of image-text fusion on the dataset
[0077]
[0078] Optional embodiments of the present application also verified the influence of different fusion methods on the recognition results in the image-text sharing representation space, as shown in Table 3. The results show that weighted fusion has the highest accuracy compared to other fusion methods.
[0079] Table 3 Performance of different fusion methods on the dataset
[0080]
[0081] For the small sample problem in the dataset, optional embodiments of the present application limited the maximum number of samples to 50 as the small sample dataset, and statistically calculated the recognition accuracy of the model on the small sample dataset. The metrics are shown in Table 4.
[0082] Table 4 Performance on the small sample dataset
[0083] Few-shot Dataset Accuracy +MobileNetv3 71.29 +InceptionV3 71.71 +ResNet50 59.62 +VGG16 50.24 +VGG19 55.46 +Xception 68.99 +DenseNet121 63.50 +NASNet 70.28 +MobileNetV2 57.39 +InceptionResNetV2 67.26 Multi-modal Fusion Scheme (Ours) 81.39
[0084] Based on the above analysis and screening process, optional embodiments of the present application determined the multi-modal fusion scheme for the wash label image text that needs to be adopted. Specifically, YoloV5 + InceptionV3 was used as the image branch feature extraction network, and OCR + IF-IDF + BERT + LSTM was used as the text branch feature extraction network. A weighted fusion method was adopted in the intermediate stage, achieving the best performance. The recognition accuracy was improved by 7.29% compared to single-modal recognition, and the best performance was also achieved on the small sample dataset, with the recognition accuracy increased by 9.68%.
[0085] In summary, optional embodiments of the present application proposed a method for identifying clothing washing and care methods; aiming at problems such as damage and reflection of clothing wash labels, provided an end-to-end multi-task learning architecture based on multi-modal fusion, utilized the method of multi-modal fusion in the image space and text space, and adopted a scheme combining image detection and classification to identify wash labels. Experiments proved that the proposed scheme has a leading advantage compared to single-modal recognition, with a 7.29% performance improvement. The present invention also showed obvious advantages on the small sample dataset, with a 9.68% increase in recognition accuracy.
[0086] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), including several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of various embodiments of the present application.
[0087] In this embodiment, a device for determining a washing and care method based on multimodal fusion is further provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated here. As used hereinafter, the term "module" can be a combination of software and / or hardware that realizes a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0088] Figure 9 It is a structural block diagram of an optional device for determining a washing and care method based on multimodal fusion according to an embodiment of the present application. The device includes:
[0089] An input module 92, configured to input a first image to be recognized into a target fusion model; wherein, the target fusion model is trained through a first data set, and the first data set includes: image data obtained by performing icon and text annotation on a second image according to a preset annotation method;
[0090] An extraction module 94, configured to extract text features corresponding to the first image through a first model branch of the target fusion model, and extract image features corresponding to the first image through a second model branch of the target fusion model;
[0091] A fusion module 96, configured to perform a feature fusion operation on the text features and the image features according to a preset fusion method to obtain a recognition result of the first image, and determine the recognition result as the washing and care method of the target clothing, where the target clothing carries a care label corresponding to the first image.
[0092] Through the above device, a first image to be recognized is input into a target fusion model; wherein, the target fusion model is trained through a first data set, and the first data set includes: image data obtained by performing icon and text annotation on a second image according to a preset annotation method; text features corresponding to the first image are extracted through a first model branch of the target fusion model, and image features corresponding to the first image are extracted through a second model branch of the target fusion model; a feature fusion operation is performed on the text features and the image features according to a preset fusion method to obtain a recognition result of the first image, and the recognition result is determined as the washing and care method of the target clothing, where the target clothing carries a care label corresponding to the first image. By adopting the above technical solution, the problem in the related art that due to the use of a single-mode architecture for the recognition of care labels and a large amount of irrelevant information on the care labels, the recognition effect of the care labels is poor is solved; thus, the recognition effect of the care label is improved.
[0093] In an exemplary embodiment, the above-mentioned device further includes a labeling module, which is used to perform a hit operation on the icons and text in the second image to obtain the hit icons and the hit text: determine the first label corresponding to the hit icons in a pre-set multi-level label system, and label the hit icons with the description text indicated by the first label to obtain the icon labeling data of the second image; and identify the first text information of the hit text according to a preset recognition method, and label the hit text with the first text information to obtain the text labeling data of the second image; wherein, the first label is the label with the highest similarity to the hit icons in the multi-level label system; determine the second image, the icon labeling data, and the text labeling data as the content included in the first data set.
[0094] In an exemplary embodiment, the labeling module is further used to perform a hit operation on the icons and text in the third image to obtain the icons of the hit third image and the text of the hit third image, wherein the wash label corresponding to the third image is different from the wash label corresponding to the second image: determine the second label corresponding to the icons of the hit third image in a pre-set multi-level label system, and label the icons of the hit third image with the description text indicated by the second label to obtain the icon labeling data of the third image; and identify the second text information of the text of the hit third image according to a preset recognition method, and label the text of the hit third image with the second text information to obtain the text labeling data of the third image; wherein, the second label is the label with the highest similarity to the text of the hit third image in the multi-level label system; determine the third image, the icon labeling data of the third image, and the text labeling data of the third image as the content included in the second data set; wherein, the second data set is used to determine the model architecture of the target fusion model.
[0095] In an exemplary embodiment, the above-mentioned device further includes a determination module, which is used to determine the label detection network allowed to be adopted by the second model branch, and determine multiple feature extraction networks allowed to be used for feature extraction of the labels detected by the label detection network; input the second data set into multiple simulation branches respectively to obtain the first precision of the multiple simulation branches for feature extraction of the second data set respectively, wherein the multiple simulation branches are simulation branches constructed by the label detection network and the multiple feature extraction networks respectively; determine the simulation branch corresponding to the highest first precision among the multiple first precisions as the second model branch.
[0096] In an exemplary embodiment, the determining module is further configured to, when the second model branch has been determined and M preset networks have been preset for the first model branch, determine a preset recognition order of the M preset networks for the first graph in the first model branch, where M is a positive integer; perform ablation experiments on the M preset networks in reverse order of the preset recognition order to obtain M experimental models corresponding to the target fusion model; input the second data set into the M experimental models respectively to determine second accuracies of the M experimental models for performing image recognition on the second data set; and determine the experimental model corresponding to the highest second accuracy among the multiple second accuracies as the target fusion model.
[0097] In an exemplary embodiment, the determining module is further configured to determine multiple feature fusion methods preset for a multi-modal fusion space included in the target fusion model, where the multi-modal fusion space is used to perform feature fusion on text features extracted by a first model branch included in the target fusion model and image features extracted by a second model branch; input the second data set into the target fusion model using different feature fusion methods respectively to determine third accuracies of the target fusion model using different feature fusion methods for performing image recognition, where the different feature fusion methods are different feature fusion methods included in the multiple feature fusion methods; and determine the feature fusion method corresponding to the highest third accuracy among the multiple third accuracies as the preset fusion method.
[0098] In an exemplary embodiment, the extraction module 94 is further configured to, when it has been determined that the first model branch includes one or a combination of a character recognition network, a text encoding network, and a recurrent neural network, convert the text image in the first image into editable text through the character recognition network; encode the editable text into a corresponding text sequence through the text encoding network; and model the text sequence through the recurrent neural network to extract text features in the text sequence.
[0099] In an exemplary embodiment, the fusion module 96 is further configured to, when the preset fusion method is weighted fusion, determine a first clarity degree of the icon in the first image and a second clarity degree of the text in the first image; determine a first fusion weight corresponding to the image features and a second fusion weight corresponding to the text features according to a proportional relationship between the first clarity degree and the second clarity degree; and determine the recognition result through the image features, the text features, the first fusion weight, and the second fusion weight.
[0100] Embodiments of the present application further provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when running.
[0101] Optionally, in this embodiment, the above storage medium may be configured to store a computer program for executing the following steps:
[0102] S1. Input a first image to be recognized into a target fusion model; wherein, the target fusion model is trained by a first data set, and the first data set includes: image data obtained by performing icon and text annotation on a second image according to a preset annotation method;
[0103] S2. Extract text features corresponding to the first image through a first model branch of the target fusion model, and extract image features corresponding to the first image through a second model branch of the target fusion model;
[0104] S3. Perform a feature fusion operation on the text features and the image features according to a preset fusion method to obtain a recognition result of the first image, and determine the recognition result as the washing and care method of the target clothing, wherein the target clothing carries a washing label corresponding to the first image.
[0105] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as a USB flash drive, a read-only memory (ROM for short), a random access memory (RAM for short), a mobile hard disk, a magnetic disk, or an optical disc that can store a computer program.
[0106] Specific examples in this embodiment may refer to the examples described in the above embodiments and exemplary embodiments, and will not be repeated here.
[0107] Embodiments of the present application further provide an electronic device including a memory and a processor, where the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above method embodiments.
[0108] Optionally, in this embodiment, the above processor may be configured to execute the following steps through a computer program:
[0109] S1. Input a first image to be recognized into a target fusion model; wherein, the target fusion model is trained by a first data set, and the first data set includes: image data obtained by performing icon and text annotation on a second image according to a preset annotation method;
[0110] S2, extract the text features corresponding to the first image through the first model branch of the target fusion model, and extract the image features corresponding to the first image through the second model branch of the target fusion model;
[0111] S3, perform a feature fusion operation on the text features and the image features according to a preset fusion method to obtain the recognition result of the first image, and determine the recognition result as the washing and care method of the target clothing, where the target clothing carries a washing label corresponding to the first image.
[0112] In an exemplary embodiment, the above electronic device may further include a transmission device and an input / output device, where the transmission device is connected to the above processor, and the input / output device is connected to the above processor.
[0113] Specific examples in this embodiment may refer to the examples described in the above embodiments and exemplary embodiments, and will not be repeated here.
[0114] Obviously, those skilled in the art should understand that the above modules or steps of the present application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. In this way, the present application is not limited to any specific combination of hardware and software.
[0115] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A method for determining a washing and care method based on multimodal fusion, characterized in that, it includes: Input the first image to be recognized into the target fusion model; wherein, the target fusion model is trained by a first data set, and the first data set includes: image data obtained by performing icon and text annotation on a second image according to a preset annotation method; Extract the text features corresponding to the first image through the first model branch of the target fusion model, and extract the image features corresponding to the first image through the second model branch of the target fusion model; Perform a feature fusion operation on the text features and the image features according to a preset fusion method to obtain the recognition result of the first image, and determine the recognition result as the washing and care method of the target clothing, wherein the target clothing carries the washing label corresponding to the first image.
2. The method for determining a washing and care method based on multimodal fusion according to claim 1, characterized in that, Before inputting the first image to be recognized into the target fusion model, the method further includes: Perform a hit operation on the icons and text in the second image to obtain the hit icons and hit text: Determine the first label corresponding to the hit icon in a preset multi-level label system, and annotate the hit icon with the description text indicated by the first label to obtain the icon annotation data of the second image; and identify the first text information of the hit text according to a preset recognition method, and annotate the hit text with the first text information to obtain the text annotation data of the second image; wherein, the first label is the label with the highest similarity to the hit icon in the multi-level label system; Determine the second image, the icon annotation data and the text annotation data as the content included in the first data set.
3. The method for determining a washing and care method based on multimodal fusion according to claim 1, characterized in that, Before inputting the first image to be recognized into the target fusion model, the method further includes: Perform a hit operation on the icons and text in the third image to obtain the hit icons of the third image and the hit text of the third image, wherein the washing label corresponding to the third image is different from the washing label corresponding to the second image: Determine the second label corresponding to the hit icon of the third image in a preset multi-level label system, and annotate the hit icon of the third image with the description text indicated by the second label to obtain the icon annotation data of the third image; and identify the second text information of the hit text of the third image according to a preset recognition method, and annotate the hit text of the third image with the second text information to obtain the text annotation data of the third image; wherein, the second label is the label with the highest similarity to the hit text of the third image in the multi-level label system; Determine the third image, the icon annotation data of the third image, and the text annotation data of the third image as the content included in the second data set; wherein, the second data set is used to determine the model architecture of the target fusion model.
4. The method for determining a washing and care method based on multimodal fusion according to claim 3, wherein, after annotating the icons in the hit third image and the text in the third image in the following manner to obtain the second data set, the method further includes: Determine the label detection network allowed to be adopted by the second model branch, and determine multiple feature extraction networks allowed to be used for feature extraction of the labels detected by the label detection network; Input the second data set into multiple simulation branches respectively to obtain the first precision of each of the multiple simulation branches for feature extraction of the second data set; wherein, the multiple simulation branches are simulation branches constructed by respectively combining the label detection network with the multiple feature extraction networks; Determine the simulation branch corresponding to the highest first precision among the multiple first precisions as the second model branch.
5. The method for determining a washing and care method based on multimodal fusion according to claim 3, wherein, after annotating the icons in the hit third image and the text in the third image in the following manner to obtain the second data set, the method further includes: When the second model branch has been determined and M preset networks have been preset for the first model branch, determine the preset recognition order of the M preset networks for the first graph in the first model branch, where M is a positive integer; Perform ablation experiments on the M preset networks respectively in the reverse order of the preset recognition order to obtain M experimental models corresponding to the target fusion model; Input the second data set into the M experimental models respectively to determine the second precision of each of the M experimental models for image recognition of the second data set; Determine the experimental model corresponding to the highest second precision among the multiple second precisions as the target fusion model.
6. The method for determining a washing and care method based on multimodal fusion according to claim 3, wherein, after annotating the icons in the hit third image and the text in the third image in the following manner to obtain the second data set, the method further includes: Determine multiple feature fusion methods preset for the multimodal fusion space included in the target fusion model, wherein the multimodal fusion space is used to perform feature fusion on the text features extracted by the first model branch included in the target fusion model and the image features extracted by the second model branch; Input the second data set into the target fusion models adopting different feature fusion methods respectively to determine the third precision of the target fusion models adopting different feature fusion methods for image recognition, wherein the different feature fusion methods are different feature fusion methods included in the multiple feature fusion methods; Determine the feature fusion method corresponding to the highest third precision among the multiple third precisions as the preset fusion method.
7. The method for determining a washing and care method based on multimodal fusion according to claim 1, wherein, extracting the text features corresponding to the first image through the first model branch of the target fusion model includes: when it is determined that the first model branch includes one or a combination of a character recognition network, a text encoding network, and a recurrent neural network, converting the text image in the first image into editable text through the character recognition network; encoding the editable text into a corresponding text sequence through the text encoding network; modeling the text sequence through the recurrent neural network to extract the text features in the text sequence.
8. The method for determining a washing and care method based on multimodal fusion according to claim 1, wherein, performing a feature fusion operation on the text features and the image features according to a preset fusion method to obtain an identification result for the first image, including: when the preset fusion method is weighted fusion, determining a first clarity degree of the icon in the first image and determining a second clarity degree of the text in the first image; determining a first fusion weight corresponding to the image features and a second fusion weight corresponding to the text features according to the proportional relationship between the first clarity degree and the second clarity degree; determining the identification result through the image features, the text features, the first fusion weight, and the second fusion weight.
9. A computer-readable storage medium, wherein, the computer-readable storage medium includes a stored program, wherein the program, when running, executes the method according to any one of claims 1 to 8.
10. An electronic device, comprising a memory and a processor, wherein, a computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 8 through the computer program.