Feature Recognition Model Training Method, Article Feature Recognition Method and Device
By extracting image features, text features and style features in the object sample data set, and adjusting them using the feature combination layer and matching layer, the feature recognition model is trained to obtain, which solves the problems of low efficiency and poor generalization of item style labeling in the prior art, and achieves more efficient item style recognition.
Patent Information
- Application Number
- CN202111585717.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-22
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-12-22
AI Technical Summary
In the prior art, item style labeling relies on manual editing rules, resulting in labor consumption, low labeling efficiency, poor generalization, and machine learning algorithms that are not generalized to unfamiliar items and styles.
By obtaining the item sample data set, including item image data, text data and style data, the initial model extracts image features, text features and style features, and inputs them into the feature combination layer and matching layer, adjusting the model parameters according to the matching results and style labels, and training obtains the feature recognition model.
The generalization of the recognition of unfamiliar creatures and unfamiliar styles has been improved, the problem of poor generalization in the existing technology has been overcome, and more efficient object style labeling and recognition has been achieved.
Smart Images

Figure CN114332477B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and more particularly, to a method for training a feature recognition model, an article feature recognition method and apparatus, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] The traditional method for item style annotation mainly uses manual annotation. This method requires a large number of manually edited rules to be set in advance, and has problems such as consuming manpower and material resources, low annotation efficiency, and poor generality. Moreover, during the process of setting the manually edited rules, due to different subjective cognitions of people, there will be differences in the setting of the editing rules, which will further lead to differences in the results of manual annotation styles.
[0003] In related technologies, machine learning algorithms are used to implement the annotation of item styles, but such machine learning algorithms have poor generalization ability for unfamiliar items and unfamiliar styles. Summary of the Invention
[0004] In view of this, the present disclosure provides a method for training a feature recognition model, an article feature recognition method and apparatus, an electronic device, a computer-readable storage medium, and a computer program product.
[0005] The first aspect of the present disclosure provides a method for training a feature recognition model, including:
[0006] Obtain an item sample data set, where the item sample data set includes multiple item samples, and each item sample includes item image data, text data for describing the item, and item style data, and where the item sample has a style label;
[0007] For each item sample, input the item image data into the image feature extraction layer of the initial model to output item image features; input the text data for describing the item into the text feature extraction layer of the initial model to output text features; input the item style data into the style feature extraction layer of the initial model to output style features;
[0008] Input the item image features and text features into the feature combination layer of the initial model to output item combination features;
[0009] Input the item combination features and style features into the matching layer of the initial model to output a matching result representing the matching of the item combination features and style features;
[0010] Adjust the model parameters of the initial model according to the matching result and the style label to obtain a trained feature recognition model.
[0011] According to an embodiment of the present disclosure, inputting the item image features and text features into the image combination layer of the initial model to output item combination features includes:
[0012] Use an image combination layer to splice the item image features and text features into item combination features.
[0013] According to an embodiment of the present disclosure, after obtaining the item sample data set, it further includes:
[0014] Generate an augmented item sample data set according to the item sample data set.
[0015] According to an embodiment of the present disclosure, generating an augmented item sample data set according to the item sample data set includes:
[0016] According to the item image data, determine a list of first item data similar to the item image data from the item database, where the list of first item data includes item image data of different items, text data for describing the items, and item style data;
[0017] According to the text data for describing the items, determine a list of second item data similar to the text data for describing the items from the item database, where the list of second item data includes item image data of different items, text data for describing the items, and item style data;
[0018] Generate an augmented item sample data set according to the list of first item data and the list of second item data.
[0019] According to an embodiment of the present disclosure, the model parameters of the initial model include the model parameters of the image feature extraction layer of the initial model, the model parameters of the text feature extraction layer of the initial model, and the model parameters of the style feature extraction layer of the initial model. Adjust the model parameters of the initial model according to the matching result and the style label to obtain a trained feature recognition model, including:
[0020] Adjust the model parameters of the image feature extraction layer of the initial model, the model parameters of the text feature extraction layer of the initial model, and the model parameters of the style feature extraction layer of the initial model according to the matching result and the style label to obtain a trained feature recognition model.
[0021] The second aspect of the present disclosure provides an item feature recognition method, including:
[0022] Obtain the item information of the item to be processed, where the item information includes item image information and text information for describing the item;
[0023] Input the item image information into the image feature extraction layer of the feature recognition model, and output the item image features of the item to be processed, where the feature recognition model is trained by the training method of the embodiment of the present disclosure;
[0024] Input the text information for describing an item into the text feature extraction layer of the feature recognition model, and output the text features of the item to be processed;
[0025] Input the item image features and text features into the feature combination layer of the feature recognition model, and output the item combination features of the item to be processed;
[0026] Input the item combination features and the description features of the candidate style into the matching layer of the feature recognition model, and output the matching result for characterizing the item combination features and the style features. Among them, the description features of the candidate style are obtained after inputting the item style information retrieved from the candidate style database into the style feature extraction layer of the feature recognition model;
[0027] Determine the style feature information that matches the item combination features of the item to be processed according to the matching result.
[0028] According to the embodiments of the present disclosure, the above feature recognition method further includes:
[0029] Obtain multiple item review text data and multiple item title text data;
[0030] Preprocess the multiple item review text data and the multiple item title text data to obtain a text data set for characterizing the item style;
[0031] Generate a candidate style database according to the text data set for characterizing the item style.
[0032] According to the embodiments of the present disclosure, the above feature recognition method further includes:
[0033] Generate a vector data set for characterizing the item style according to the text data set for characterizing the item style;
[0034] Generate a candidate style database according to the vector data set for characterizing the item style.
[0035] The third aspect of the present disclosure provides a feature recognition model training device, including: a first acquisition module, a feature extraction module, a feature combination module, a matching module, and an adjustment module. Among them, the first acquisition module is configured to acquire an item sample data set, where the item sample data set includes multiple item samples, and each item sample includes item image data, text data for describing the item, and item style data, and the item sample has a style label. The feature extraction module is configured to, for each item sample, input the item image data into the image feature extraction layer of the initial model and output the item image feature; input the text data for describing the item into the text feature extraction layer of the initial model and output the text feature; input the item style data into the style feature extraction layer of the initial model and output the style feature. The feature combination module is configured to input the item image feature and the text feature into the feature combination layer of the initial model and output the item combination feature. The matching module is configured to input the item combination feature and the style feature into the matching layer of the initial model and output a matching result for characterizing the item combination feature and the style feature. The adjustment module is configured to adjust the model parameters of the initial model according to the matching result and the style label to obtain a trained feature recognition model.
[0036] According to an embodiment of the present disclosure, the feature combination module includes a splicing unit configured to splice the item image feature and the text feature into the item combination feature by using an image combination layer.
[0037] According to an embodiment of the present disclosure, the above device further includes: a generation module configured to generate an augmented item sample data set according to the item sample data set.
[0038] According to an embodiment of the present disclosure, the generation module includes a first determination unit, a second determination unit, and a generation unit. Among them, the first determination unit is configured to determine a first list of item data similar to the item image data from an item database according to the item image data, where the first list of item data includes item image data of different items, text data for describing the items, and item style data. The second determination unit is configured to determine a second list of item data similar to the text data for describing the item from the item database according to the text data for describing the item, where the second list of item data includes item image data of different items, text data for describing the items, and item style data. The generation unit is configured to generate an augmented item sample data set according to the first list of item data and the second list of item data.
[0039] According to an embodiment of the present disclosure, the generation unit includes a generation subunit configured to generate an augmented item sample data set according to the data intersection of the first list of item data and the second list of item data.
[0040] According to an embodiment of the present disclosure, the adjustment module includes an adjustment unit configured to adjust model parameters of the image feature extraction layer of the initial model, model parameters of the text feature extraction layer of the initial model, and model parameters of the style feature extraction layer of the initial model according to the matching result and the style label, so as to obtain a trained feature recognition model.
[0041] A fourth aspect of the present disclosure provides an article feature recognition device, including: a second acquisition module, a first recognition module, a second recognition module, a third recognition module, a fourth recognition module, and a determination module. Among them, the second acquisition module is configured to acquire article information of an article to be processed, where the article information includes article image information and text information for describing the article. The first recognition module is configured to input the article image information into the image feature extraction layer of the feature recognition model and output the article image features of the article to be processed, where the feature recognition model is trained by the training method provided by the embodiment of the present disclosure. The second recognition module is configured to input the text information for describing the article into the text feature extraction layer of the feature recognition model and output the text features of the article to be processed. The third recognition module is configured to input the article image features and the text features into the feature combination layer of the feature recognition model and output the article combination features of the article to be processed. The fourth recognition module is configured to input the article combination features and the description features of the candidate style into the matching layer of the feature recognition model and output a matching result for characterizing the matching between the article combination features and the description features of the candidate style, where the description features of the candidate style are obtained after inputting the article style information acquired from the candidate style database into the style feature extraction layer of the feature recognition model. The determination module is configured to determine style feature information that matches the article combination features of the article to be processed according to the matching result.
[0042] A fifth aspect of the present disclosure provides an electronic device, including: one or more processors; a memory configured to store one or more programs, where when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above-mentioned feature recognition model training method.
[0043] A sixth aspect of the present disclosure further provides a computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor is caused to execute the above-mentioned feature recognition model training method.
[0044] A seventh aspect of the present disclosure further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the above-mentioned feature recognition model training method is implemented.
[0045] According to an embodiment of the present disclosure, because an item sample data set including item image data, text data for describing the item, item style data, and style labels is adopted, by respectively inputting the item image data and the text data for describing the item into the image feature extraction layer and the text feature extraction layer of the initial model to output image features and text features, and then inputting them into the feature combination layer to output item combination features, inputting the item combination features and the style features extracted by the style feature extraction layer of the initial model into the matching layer, and training to obtain a feature recognition model by adjusting the model parameters of the initial model through the output matching result and the style labels, at least partially overcoming the technical problem of poor generalization of unfamiliar items and styles in the related art by using machine learning algorithms, and improving the generalization of the recognition of unfamiliar items and unfamiliar styles. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, the above and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:
[0047] Figure 1 Schematically shows an exemplary system architecture to which the feature recognition model training method according to the embodiments of the present disclosure can be applied;
[0048] Figure 2 Schematically shows a flowchart of the feature recognition model training method according to the embodiments of the present disclosure;
[0049] Figure 3 Schematically shows a flowchart of the method for amplifying the sample data set according to the embodiments of the present disclosure;
[0050] Figure 4 Schematically shows a schematic diagram of the feature recognition model architecture according to the embodiments of the present disclosure;
[0051] Figure 5 Schematically shows a flowchart of the item feature recognition method according to the embodiments of the present disclosure;
[0052] Figure 6 Schematically shows a flowchart of the method for generating a candidate style database according to the embodiments of the present disclosure;
[0053] Figure 7 Schematically shows a system architecture diagram of the item feature recognition according to the embodiments of the present disclosure;
[0054] Figure 8 Schematically shows a block diagram of the feature recognition model training device according to the embodiments of the present disclosure;
[0055] Figure 9 Schematically shows a block diagram of the item feature recognition device according to the embodiments of the present disclosure; and
[0056] Figure 10 FIG. 1 schematically shows a block diagram of an electronic device suitable for implementing a feature recognition model training method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, numerous specific details are set forth in order to provide a thorough understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments may be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present disclosure.
[0058] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0059] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0060] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning commonly understood by those of ordinary skill in the art (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C).
[0061] Embodiments of the present disclosure provide a feature recognition model training method, which uses an item sample data set including item image data, text data for describing the item, item style data, and item style labels. By inputting the item image data and the text data for describing the item into the image feature extraction layer and the text feature extraction layer of the initial model respectively to output image features and text features, and then inputting them into the feature combination layer to output item combination features, and inputting the item combination features and the style features extracted by the style feature extraction layer of the initial model into the matching layer, the model parameters of the initial model are adjusted by the output matching result and the style label, and a feature recognition model is trained by means of the above technical means.
[0062] Figure 1Schematically shown is an exemplary system architecture 100 to which the feature recognition model training method according to an embodiment of the present disclosure can be applied. It should be noted that Figure 1 What is shown is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.
[0063] As Figure 1 shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0064] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only as examples).
[0065] The terminal devices 101, 102, 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0066] The server 105 may be a server providing various services, such as a background management server that supports the websites browsed by users using the terminal devices 101, 102, 103 (only as an example). The background management server may analyze and process data such as received user requests, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0067] It should be noted that the feature recognition model training method or feature recognition method provided by the embodiments of the present disclosure can generally be executed by the server 105. Correspondingly, the feature recognition model training device or feature recognition device provided by the embodiments of the present disclosure can generally be set in the server 105. The feature recognition model training method or feature recognition method provided by the embodiments of the present disclosure can also be executed by a server or server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103, and / or the server 105. Correspondingly, the feature recognition model training device or feature recognition device provided by the embodiments of the present disclosure can also be set in a server or server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103, and / or the server 105. Alternatively, the feature recognition model training method or feature recognition method provided by the embodiments of the present disclosure can also be executed by the terminal devices 101, 102, or 103, or can also be executed by other terminal devices different from the terminal devices 101, 102, or 103. Correspondingly, the feature recognition model training device or feature recognition device provided by the embodiments of the present disclosure can also be set in the terminal devices 101, 102, or 103, or set in other terminal devices different from the terminal devices 101, 102, or 103.
[0068] For example, the item sample data set can be obtained through any one of the terminal devices 101, 102, or 103 (for example, the terminal device 101, but not limited thereto), can also be stored in any one of the terminal devices 101, 102, or 103, or stored on an external storage device and can be imported into the terminal device 101. Then, the terminal device 101 can execute the feature recognition model training method provided by the embodiments of the present disclosure locally, or send the item sample data set to other terminal devices, servers, or server clusters, and the other terminal devices, servers, or server clusters that receive the item sample data set execute the feature recognition model training method provided by the embodiments of the present disclosure.
[0069] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in
[0070] Figure 2 is merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.
[0071] As Figure 2 shown, the feature recognition model training method of this embodiment includes operations S210 to S270.
[0072] In operation S210, an item sample dataset is obtained. The item sample dataset includes multiple item samples, and each item sample includes item image data, text data for describing the item, and item style data. Here, the item samples have style labels.
[0073] According to an embodiment of the present disclosure, the item image data may include an item picture. In addition to including the item itself, the item picture may also include a model or a display stand, a display cabinet, etc. for displaying the item. The text data for describing the item may include text data for describing the color of the item, text data for describing the brand of the item, text data for describing the style of the item, etc. For example: a red pleated skirt, a blue T-shirt of a certain sports brand, etc. The item style data may include a sports style, a casual style, a business style, a sweet style, etc. The style label may use a binary number to indicate whether the item matches the style. When the item matches the style, the style label is 1; when the item does not match the style, the style label is 0. For example: if the style of a blue T-shirt of a certain sports brand is a sports style, then the style label is 1. If the style of a blue T-shirt of a certain sports brand is a ladylike style, then the style label is 0.
[0074] In operation S220, for each item sample, the item image data is input into the image feature extraction layer of the initial model, and item image features are output.
[0075] According to an embodiment of the present disclosure, for example: the item image data is a picture of a blue T-shirt of a certain sports brand. When the item image data is input into the image feature extraction layer of the initial model, the output item image features may include blue, the logo of the sports brand, the T-shirt, etc.
[0076] In operation S230, the text data for describing the item is input into the text feature extraction layer of the initial model, and text features are output.
[0077] According to an embodiment of the present disclosure, for example: the text data for describing the item is a blue T-shirt of a certain sports brand. When the item image data is input into the text data extraction layer of the initial model, the output text data features may include blue, the name of the sports brand, the T-shirt, etc.
[0078] In operation S240, the item style data is input into the style feature extraction layer of the initial model, and style features are output.
[0079] According to an embodiment of the present disclosure, for example: the style of the blue T-shirt of a certain sports brand is a sports style, then the item style label is 1, indicating that the item matches the style, and the sample data is positive sample data. When the item style data is input into the style feature extraction layer of the initial model, the output style features may be sports.
[0080] In operation S250, the item image features and text features are input into the feature combination layer of the initial model, and item combined features are output.
[0081] According to an embodiment of the present disclosure, the image features of the above item, i.e., "blue, logo of a sports brand, T-shirt", and the text features, i.e., "blue, name of a sports brand, T-shirt", are input into the feature combination layer, and the output item combined features are "blue, logo of a sports brand, name of a sports brand, T-shirt".
[0082] In operation S260, the item combined features and style features are input into the matching layer of the initial model, and a matching result representing the item combined features and style features is output.
[0083] According to an embodiment of the present disclosure, the item combined features "blue, logo of a sports brand, name of a sports brand, T-shirt" and the style feature "sports" are input into the matching layer of the initial model, and a matching result representing the item combined features and style features is output. For example, the matching result is 0.8.
[0084] According to an embodiment of the present disclosure, the item combined features and style features can be represented by vectors, and the matching layer of the initial model can determine the matching result representing the item combined features and style features by calculating the distance between the item combined feature vector and the style feature vector, such as Euclidean distance, cosine similarity, etc.
[0085] In operation S270, the model parameters of the initial model are adjusted according to the matching result and the style label, and a trained feature recognition model is obtained.
[0086] According to an embodiment of the present disclosure, since the above input is positive sample data and the style label is 1, therefore, the model parameters of the initial model can be adjusted according to the matching result 0.8 and the style label 1. After the adjustment is completed, the output matching result can become 0.9, 0.95 or a value closer to 1. The closer the output matching result is to 1, the higher the recognition accuracy of the feature recognition model is proved.
[0087] According to an embodiment of the present disclosure, because an item sample data set including item image data, text data for describing the item, item style data, and item style labels is adopted, the item image data and the text data for describing the item are respectively input into the image feature extraction layer and the text feature extraction layer of the initial model to output image features and text features, and then input into the feature combination layer to output item combination features. The item combination features and the style features extracted by the style feature extraction layer of the initial model are input into the matching layer, and the model parameters of the initial model are adjusted through the output matching results and style labels. By using this technical means to train the feature recognition model, at least partially, the technical problem of poor generalization of unfamiliar items and styles in the related art using machine learning algorithms is overcome, and the generalization of the recognition of unfamiliar items and unfamiliar styles is improved.
[0088] According to an embodiment of the present disclosure, inputting the item image features and text features into the image combination layer of the initial model, the output item combination features include:
[0089] Using the image combination layer to splice the item image features and text features into item combination features.
[0090] According to an embodiment of the present disclosure, both the item image features and text features can be represented by vectors. Using the image combination layer, the item image feature vector and the text feature vector can be spliced into a one-dimensional vector. For example: the item image feature vector is (a, b, c), and the text feature vector is (e, f, g), then the spliced item combination feature vector can be (a, b, c, e, f, g).
[0091] According to an embodiment of the present disclosure, the item image feature vector and the text feature vector can also be spliced into a two-dimensional vector or matrix using the image combination layer, which will not be elaborated here.
[0092] According to an embodiment of the present disclosure, by obtaining the item combination features through splicing, the image features and text features of the item can be combined to identify the style matching the item, improving the accuracy of model prediction.
[0093] According to an embodiment of the present disclosure, after obtaining the item sample data set, the above training method further includes:
[0094] Generating an augmented item sample data set according to the item sample data set.
[0095] According to an embodiment of the present disclosure, the augmented item sample data set includes, on the basis of the obtained item sample data set, respectively performing deep representation on the picture data and text data in the item sample data set through two methods of picture data augmentation and text data augmentation to obtain the augmented item sample data set.
[0096] According to an embodiment of the present disclosure, for example, the item image data included in the item sample data may be a picture of black sports shoes with a certain brand logo, the text data for describing the item may be a certain brand of black sports shoes, the item style data is a sports style, and the item style label is 1. According to the common features of the item sample data, the item sample data can be amplified. For example, the amplified item sample data may be the item image data of red or green sports shoes with a certain brand logo, the text data for describing the item, the item style data, and the item style label, or may be the item image data of various different colored sports shoes without a certain brand logo and with a style similar to that in the item sample data, the text data for describing the item, the item style data, and the style label, etc. The amplified item sample data listed above for the item sample data set are all positive sample data with the item and the style matching, so the style label is 1.
[0097] According to an embodiment of the present disclosure, when performing model training, negative sample data with the item and the style not matching can also be used. For example, the item image data included in the negative sample data may be a picture of black sports shoes with a certain brand logo, the text data for describing the item may be a certain brand of black sports shoes, the item style data is a business style, and the style label is 0.
[0098] According to an embodiment of the present disclosure, when implementing the feature recognition model training method of the embodiments of the present disclosure, the ratio range of the positive sample data to the negative sample data can be between 1:1 and 1:5.
[0099] According to an embodiment of the present disclosure, since the item sample data set mainly comes from domain expert data and a small amount of manually labeled data, the data stream is limited. Generating an amplified item sample data set and using the amplified item sample data set and the item sample data to train the initial model can improve the accuracy of the feature recognition model in recognition and prediction.
[0100] The following refers to Figures 3 to 7 and further describes the method shown in Figure 2 in combination with specific embodiments.
[0101] Figure 3 Schematically shows a flowchart of a method for generating an amplified sample data set according to an embodiment of the present disclosure.
[0102] As Figure 3 shown, the generation of the amplified sample data set in this embodiment includes operations S310 to S330.
[0103] In operation S310, according to the item image data, a list of first item data similar to the item image data is determined from the item database, where the list of first item data includes the item image data of different items, the text data for describing the item, and the item style data.
[0104] According to an embodiment of the present disclosure, for example, taking the item image data as a picture of a pink dress with lace as an example, a list of first item data similar to the item image data can be determined from the item database, and the list can include multiple different items, such as: a red dress, a skirt with lace, and a pink dress with polka dots. Taking the red dress item in the list of first item data as an example, the list of first item data includes a picture of the red dress, such as: a picture of a little girl wearing the red dress; text data for describing the red dress, such as: red, dress; item style data, such as: sweet style, ladylike style.
[0105] In operation S320, according to the text data for describing the item, a list of second item data similar to the text data for describing the item is determined from the item database, where the list of second item data includes item image data of different items, text data for describing the item, and item style data.
[0106] According to an embodiment of the present disclosure, for example, taking the text data of a pink dress with lace as an example, from the item database, a list of second item data similar to the text data for describing the item can be determined, and the list includes multiple different items, such as: a skirt with lace, a red dress with lace, and a pink dress with polka dots. Taking the skirt with lace included in the list of second item data as an example, the item image data included in the list of second item data is a picture of a skirt with lace, the text data for describing the item is lace, skirt, and the style data of the item is ladylike style, sweet style.
[0107] In operation S330, an augmented item sample data set is generated according to the list of first item data and the list of second item data.
[0108] According to an embodiment of the present disclosure, an augmented item sample data set can be generated according to all the data in the list of first item data and the list of second item data, and the generated augmented item sample data set includes item image data, text data for describing the item, and item style data of a red dress, a skirt with lace, a pink dress with polka dots, and a red dress with lace.
[0109] According to an embodiment of the present disclosure, by generating a list of item data through item image data and text data for describing the item respectively, item sample data can be automatically augmented to solve the problems in the related art that only relying on the obtained sample data for model training results in low model prediction accuracy and poor generalization.
[0110] According to an embodiment of the present disclosure, an augmented item sample data set is generated based on a first item data list and a second item data list, including:
[0111] An augmented item sample data set is generated based on the data intersection of the first item data list and the second item data list.
[0112] According to an embodiment of the present disclosure, for example: the first item data list includes a red dress, a knee-length skirt with lace trim, and a pink dress with polka dots. The second item data list includes a knee-length skirt with lace trim, a red dress with lace trim, and a pink dress with polka dots. Then the generated augmented item sample data set includes a picture with an item image data of a pink knee-length skirt with lace trim, and the text data for describing the item includes: lace trim, pink, dress, knee-length skirt, red, and the style of the item includes: ladylike style, sweet style.
[0113] According to an embodiment of the present disclosure, item data lists are respectively generated through item image data and text data for describing the item, and then the intersections of the two data lists are taken, which not only automatically augments the sample data but also improves the accuracy of the augmented sample data.
[0114] According to an embodiment of the present disclosure, the model parameters of the initial model include the model parameters of the image feature extraction layer of the initial model, the model parameters of the text feature extraction layer of the initial model, and the model parameters of the style feature extraction layer of the initial model. The model parameters of the initial model are adjusted according to the matching result and the style label to obtain a trained feature recognition model, including:
[0115] According to the matching result and the style label, the model parameters of the image feature extraction layer of the initial model, the model parameters of the text feature extraction layer of the initial model, and the model parameters of the style feature extraction layer of the initial model are adjusted to obtain a trained feature recognition model.
[0116] According to an embodiment of the present disclosure, the style label can be represented by a binary number. When the style label is 1, it indicates that the item in the item sample data matches the style; when the style label is 0, it indicates that the item in the item sample data does not match the style.
[0117] According to an embodiment of the present disclosure, for example, after inputting positive item sample data in which the item matches the style into the initial model, the output matching result is 0.8, and the style label is 1. Then, the model parameters of the image feature extraction layer, the text feature extraction layer, and the style feature extraction layer in the initial model can be adjusted to change the output matching result.
[0118] According to an embodiment of the present disclosure, the difference between the matching result and the actual result can be calculated using the formula shown in Equation (1).
[0119]
[0120] Among them, XOR is the exclusive OR function. When two numbers with values of 0 or 1 are equal, the return value of the XOR function is 0; otherwise, it is 1.
[0121] (Item i , Style j ) true is the matching relationship between the actual item and the style, with a value of 0 or 1. 1 indicates correspondence, and 0 indicates non - correspondence.
[0122] (Item i , Style j ) pred is the matching relationship between the item and the style predicted by the algorithm, and the value range is a floating - point number between 0 and 1. The larger the value, the higher the degree of matching between the item and the style considered by the algorithm.
[0123] I(·) is an indicator function. Here, when the input value is less than 0, it returns 0, and when the input value is greater than 0, it returns 1.
[0124] According to the embodiments of the present disclosure, multiple matching results and style labels output after multiple adjustments of the model parameters can be input into the loss function. When the change of the loss function approaches zero, at this time, the matching degrees of the image features, text features, and style features extracted by the image feature extraction layer, text feature extraction layer, and style feature extraction layer in the initial model are the highest, indicating that the prediction accuracy of the feature recognition model is the highest, and then the training of the initial model is completed to obtain the trained feature recognition model.
[0125] According to the embodiments of the present disclosure, the image feature extraction layer, text feature extraction layer, and style feature extraction layer of the initial model are adjusted through the matching results output by the matching layer and the style labels, realizing a supervised model training process and improving the prediction accuracy of model recognition.
[0126] Figure 4 Schematically shows a schematic diagram of the feature recognition model architecture according to the embodiments of the present disclosure.
[0127] As Figure 4 shown, the feature recognition model of this embodiment includes a three - layer architecture: the bottom layer includes an image feature extraction layer, a text feature extraction layer, and a style feature extraction layer, which are respectively used to extract the image features, text features, and style features in the item sample data. The middle layer includes a feature combination layer, which is used to combine the image features and text features to form item combination features. The top layer includes a matching layer, which is used to match the item combination features with the style features, output matching results, and adjust the model parameters of the bottom layer according to the matching results and style labels.
[0128] Figure 5 Schematically shows a flowchart of an article feature recognition method according to an embodiment of the present disclosure.
[0129] As Figure 5 shown, the article feature recognition method of this embodiment includes operations S510 to S560.
[0130] In operation S510, obtain the article information of the article to be processed, where the article information includes article image information and text information for describing the article.
[0131] According to an embodiment of the present disclosure, for example: the article image information included in the article information of the article to be processed is a picture of a dark blue denim jacket with a fur collar and a specific brand exclusive pattern, and the text information for describing the article is a certain brand dark blue fur collar denim jacket.
[0132] In operation S520, input the article image information into the image feature extraction layer of the feature recognition model, and output the article image features of the article to be processed, where the feature recognition model is trained by the training method of the embodiment of the present disclosure.
[0133] According to an embodiment of the present disclosure, input the picture of the dark blue denim jacket with a fur collar and a specific brand exclusive pattern into the image feature extraction layer of the feature recognition model trained by the training method provided in the embodiment of the present disclosure, and the output article image features may include: dark blue, fur collar, denim jacket, a specific brand exclusive pattern, for example, a four-leaf clover pattern.
[0134] In operation S530, input the text information for describing the article into the text feature extraction layer of the feature recognition model, and output the text features of the article to be processed.
[0135] According to an embodiment of the present disclosure, input the text information of a certain brand dark blue fur collar denim jacket into the text feature extraction layer of the feature recognition model, and the output text features of the article to be processed may include: dark blue, a certain brand, fur collar, denim jacket.
[0136] In operation S540, input the article image features and text features into the feature combination layer of the feature recognition model, and output the article combination features of the article to be processed.
[0137] According to an embodiment of the present disclosure, input the article image features "dark blue, fur collar, denim jacket, a specific brand exclusive pattern, for example, a four-leaf clover pattern" and the text features "dark blue, a certain brand, fur collar, denim jacket" into the feature combination layer of the feature recognition model, and the output article combination features of the article to be processed may be "dark blue, fur collar, denim jacket, four-leaf clover pattern, a certain brand".
[0138] In operation S550, the item combination features and the descriptive features of the candidate styles are input into the matching layer of the feature recognition model, and a matching result for characterizing the item combination features and the style features is output. Among them, the descriptive features of the candidate styles are obtained after inputting the item style information retrieved from the candidate style database into the style feature extraction layer of the feature recognition model.
[0139] According to an embodiment of the present disclosure, for example: the item combination features "dark blue, fur collar, denim jacket, four-leaf clover pattern, a certain brand" and the descriptive features of the candidate styles, for example, the descriptive features of the candidate style are casual style, are input into the matching layer of the feature recognition model, and a matching result is output.
[0140] In operation S560, according to the matching result, the style feature information that matches the item combination features of the product to be processed is determined.
[0141] According to an embodiment of the present disclosure, multiple item style information can be retrieved from the candidate style database and input into the style feature extraction layer of the feature recognition model to obtain multiple style features. The item combination features and the multiple style features are input into the matching layer of the feature recognition model, and the multiple obtained matching results are sorted. If only one style is to be retained for the item to be processed, the style with the highest matching result can be used as the item style information of the item to be processed. If N styles can be retained for the item to be processed, from the sorted matching results, the top N styles corresponding to the matching results can be taken in sequence as the item style information of the item to be processed.
[0142] According to an embodiment of the present disclosure, through the feature recognition model, the item image features and text features that match the style can be recognized, and the item image features and text features are combined to form item combination features. The style features are extracted from the candidate style database using the style feature extraction layer in the feature recognition model, and the style features that match the item combination features are determined, thereby determining the style of the item, achieving the technical effect of automatically determining the style of a new item.
[0143] Figure 6 Schematically shows a flowchart of a method for generating a candidate style database according to an embodiment of the present disclosure.
[0144] As Figure 6 shown, the method for generating a candidate style database in this embodiment includes: S610 to S630.
[0145] In operation S610, multiple item review text data and multiple item title text data are retrieved.
[0146] According to an embodiment of the present disclosure, the item review text data may include that this piece of clothing is suitable for home, travel or any casual occasion, comfortable to wear, and so on. The item title text data may include essential for home and travel, comfortable, and so on.
[0147] In operation S620, multiple item review text data and multiple item title text data are preprocessed to obtain a text data set for characterizing the item style.
[0148] According to an embodiment of the present disclosure, the preprocessing may include cleaning the data, such as word segmentation, removing stop words, removing punctuation marks, removing feature symbols, filtering words with too high frequencies, filtering words with too low frequencies, and so on. The preprocessing further includes performing clustering calculation on the cleaned data to obtain a text data set for characterizing the item style. For example: casual, sporty, and so on.
[0149] In operation S630, a candidate style database is generated according to the text data set for characterizing the item style.
[0150] According to an embodiment of the present disclosure, the text data set for characterizing the item style may be stored in a database to generate a candidate style database.
[0151] According to an embodiment of the present disclosure, by obtaining the review data and title data of an item and performing preprocessing on the data, a text data set for characterizing the item style is obtained. The problem of lacking a candidate style database or difficulty in automatically generating a candidate style database during the machine learning process is solved.
[0152] According to an embodiment of the present disclosure, the method for generating a candidate style database further includes:
[0153] Generating a vector data set for characterizing the item style according to the text data set for characterizing the item style.
[0154] Generating a candidate style database according to the vector data set for characterizing the item style.
[0155] According to an embodiment of the present disclosure, the data in the candidate style database can be represented by vectors. When matching styles for new items, style data can be queried from the candidate style database through vector retrieval, which can save matching time.
[0156] In order to compare the prediction times of vector retrieval and ordinary retrieval, the complexities of the prediction times of the two retrieval methods can be represented by Equation (2) and Equation (3).
[0157] cost=|I|*|J|*d (2)
[0158]
[0159] Among them, t is the number of iterations, |J| is the number of styles, k is the number of cluster centers, |I| is the number of items, c is the number of the nearest cluster centers to be found, and d is the vector dimension of the model training result.
[0160] Equation (2) represents the predicted time complexity of ordinary retrieval, and Equation (3) represents the predicted time complexity of brute-force retrieval.
[0161] Since both t and c are constants, and k is much smaller than |J| and also much smaller than |I|, the calculation result of Equation (3) is much smaller than that of Equation (2), that is, the predicted time of using the vector retrieval algorithm is shorter than that of the ordinary retrieval algorithm.
[0162] Figure 7 Schematically shows a system architecture diagram for item feature recognition according to an embodiment of the present disclosure.
[0163] As Figure 7 shown, the item feature recognition system includes three parts: a model training module, a candidate style database generation module, and a model prediction module.
[0164] The model training module expands the sample data by obtaining domain expert data and manually labeled data, and uses the expanded sample data to train an initial model to obtain a feature recognition model.
[0165] The candidate style database generation module obtains a candidate style database by obtaining item review data and item title data according to the method for generating a candidate style database in the embodiment of the present disclosure.
[0166] The model prediction module inputs the item image data of the item to be processed into the image feature extraction layer to extract the image feature of the item to be processed; extracts the text feature of the item to be processed from the text data for describing the item; and combines the image feature and the text feature of the item to be processed by using the item feature combination layer to obtain an item combined feature. Matches the item combined feature with the description feature of the candidate style extracted by the style feature extraction layer in the matching layer to determine the style of the item to be processed.
[0167] Figure 8 Schematically shows a block diagram of a feature recognition model training device according to an embodiment of the present disclosure.
[0168] As Figure 8 shown, the feature recognition model training device 800 includes: a first acquisition module 810, a feature extraction module 820, a feature combination module 830, a matching module 840, and an adjustment module 850.
[0169] The first acquisition module 810 is configured to acquire an item sample data set, where the item sample data set includes multiple item samples, and each item sample includes item image data, text data for describing the item, and item style data, and the item sample has an item style label.
[0170] The feature extraction module 820 is configured to, for each item sample, input the item image data into the image feature extraction layer of the initial model to output item image features; input the text data for describing the item into the text feature extraction layer of the initial model to output text features; and input the item style data into the style feature extraction layer of the initial model to output style features.
[0171] The feature combination module 830 is configured to input the item image features and text features into the feature combination layer of the initial model to output item combined features.
[0172] The matching module 840 is configured to input the item combined features and style features into the matching layer of the initial model to output a matching result for characterizing the item combined features and style features.
[0173] The adjustment module 850 is configured to adjust the model parameters of the initial model according to the matching result and the style label to obtain a trained feature recognition model.
[0174] According to an embodiment of the present disclosure, the feature combination module 830 includes a splicing unit configured to splice the item image features and text features into item combined features by using an image combination layer.
[0175] According to an embodiment of the present disclosure, the above device further includes: a generation module configured to generate an augmented item sample data set according to the item sample data set.
[0176] According to an embodiment of the present disclosure, the generation module includes a first determination unit, a second determination unit, and a generation unit. The first determination unit is configured to determine, according to the item image data, a first list of item data similar to the item image data from an item database, where the first list of item data includes item image data of different items, text data for describing the items, and item style data. The second determination unit is configured to determine, according to the text data for describing the item, a second list of item data similar to the text data for describing the item from the item database, where the second list of item data includes item image data of different items, text data for describing the items, and item style data. The generation unit is configured to generate an augmented item sample data set according to the first list of item data and the second list of item data.
[0177] According to an embodiment of the present disclosure, the generation unit includes a generation subunit configured to generate an augmented item sample data set according to the data intersection of the first list of item data and the second list of item data.
[0178] According to an embodiment of the present disclosure, the adjustment module includes an adjustment unit, configured to adjust the model parameters of the image feature extraction layer of the initial model, the model parameters of the text feature extraction layer of the initial model, and the model parameters of the style feature extraction layer of the initial model according to the matching result and the style label, so as to obtain a trained feature recognition model.
[0179] Figure 9 A block diagram of an article feature recognition device according to an embodiment of the present disclosure is schematically shown.
[0180] As Figure 9 shown, the article feature recognition device of this embodiment includes a second acquisition module 910, a first recognition module 920, a second recognition module 930, a third recognition module 940, a fourth recognition module 950, and a determination module 960.
[0181] The second acquisition module 910 is configured to acquire article information of an article to be processed, where the article information includes article image information and text information for describing the article.
[0182] The first recognition module 920 is configured to input the article image information into the image feature extraction layer of the feature recognition model, and output article image features of the article to be processed, where the feature recognition model is trained by the training method provided by the embodiment of the present disclosure.
[0183] The second recognition module 930 is configured to input the text information for describing the article into the text feature extraction layer of the feature recognition model, and output the text features of the article to be processed.
[0184] The third recognition module 940 is configured to input the article image features and the text features into the feature combination layer of the feature recognition model, and output article combination features of the article to be processed.
[0185] The fourth recognition module 950 is configured to input the article combination features and the description features of the candidate style into the matching layer of the feature recognition model, and output a matching result for characterizing the matching between the article combination features and the description features of the candidate style, where the description features of the candidate style are obtained by inputting the article style information acquired from the candidate style database into the style feature extraction layer of the feature recognition model.
[0186] The determination module 960 is configured to determine style feature information that matches the article combination features of the article to be processed according to the matching result. It should be noted that the embodiments of the device part of the present disclosure are the same as or similar to the embodiments of the method part of the present disclosure, and the present disclosure will not be elaborated herein.
[0187] Any of a plurality of modules, units, and subunits according to embodiments of the present disclosure, or at least some functions of any of them, may be implemented in one module. Any one or more of the modules, units, and subunits according to embodiments of the present disclosure may be split into multiple modules for implementation. Any one or more of the modules, units, and subunits according to embodiments of the present disclosure may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or may be implemented by hardware or firmware in any other reasonable manner of integrating or packaging circuits, or may be implemented in any one of the three implementation manners of software, hardware, and firmware, or in any appropriate combination of several of them. Alternatively, one or more of the modules, units, and subunits according to embodiments of the present disclosure may be at least partially implemented as a computer program module, which may perform corresponding functions when the computer program module is run.
[0188] For example, any combination of the first acquisition module 810, the feature extraction module 820, the feature combination module 830, the matching module 840, the adjustment module 850, or the second acquisition module 910, the first recognition module 920, the second recognition module 930, the third recognition module 940, the fourth recognition module 950, and the determination module 960 can be combined and implemented in one module / unit / sub-unit, or any one of the modules / units / sub-units can be split into multiple modules / units / sub-units. Alternatively, at least part of the functions of one or more of these modules / units / sub-units can be combined with at least part of the functions of other modules / units / sub-units and implemented in one module / unit / sub-unit. According to an embodiment of the present disclosure, at least one of the first acquisition module 810, the feature extraction module 820, the feature combination module 830, the matching module 840, the adjustment module 850, or the second acquisition module 910, the first recognition module 920, the second recognition module 930, the third recognition module 940, the fourth recognition module 950, and the determination module 960 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable means such as integrating or packaging the circuit, etc., in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the first acquisition module 810, the feature extraction module 820, the feature combination module 830, the matching module 840, the adjustment module 850, or the second acquisition module 910, the first recognition module 920, the second recognition module 930, the third recognition module 940, the fourth recognition module 950, and the determination module 960 can be at least partially implemented as a computer program module, and when the computer program module is run, it can execute the corresponding functions.
[0189] Figure 10 FIG. schematically shows a block diagram of an electronic device suitable for implementing the method described above according to an embodiment of the present disclosure. Figure 10 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0190] As Figure 10As shown, the electronic device 1000 according to an embodiment of the present disclosure includes a processor 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage section 1008 into a random access memory (RAM) 1003. The processor 1001 can include, for example, a general-purpose microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application-specific integrated circuit (ASIC)), and so on. The processor 1001 can also include on-board memory for caching purposes. The processor 1001 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0191] In the RAM 1003, various programs and data required for the operation of the electronic device 1000 are stored. The processor 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. The processor 1001 performs various operations of the method flow according to an embodiment of the present disclosure by executing the programs in the ROM 1002 and / or the RAM 1003. It should be noted that the program can also be stored in one or more memories other than the ROM 1002 and the RAM 1003. The processor 1001 can also perform various operations of the method flow according to an embodiment of the present disclosure by executing the programs stored in the one or more memories.
[0192] According to an embodiment of the present disclosure, the electronic device 1000 can further include an input / output (I / O) interface 1005, and the input / output (I / O) interface 1005 is also connected to the bus 1004. The system 1000 can further include one or more of the following components connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, etc.; an output section 1007 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN card, a modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1010 as needed so that a computer program read from it can be installed into the storage section 1008 as needed.
[0193] According to an embodiment of the present disclosure, the method flow according to the embodiment of the present disclosure can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 1009 and / or installed from the removable medium 1011. When the computer program is executed by the processor 1001, the above functions defined in the system of the embodiment of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described system, device, apparatus, module, unit, etc. can be implemented by computer program modules.
[0194] The present disclosure also provides a computer-readable storage medium, which may be included in the device / device / system described in the above embodiment; or may exist alone without being assembled into the device / device / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present disclosure is implemented.
[0195] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium. For example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, device, or device.
[0196] For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include one or more memories other than the above-described ROM 1002 and / or RAM 1003 and / or ROM 1002 and RAM 1003.
[0197] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.
[0198] Those skilled in the art will appreciate that the features recited in the various embodiments and / or claims of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly recited in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features recited in the various embodiments and / or claims of the present disclosure can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.
[0199] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present disclosure.
Claims
1. A method for training a feature recognition model, comprising: Obtain an item sample data set, where the item sample data set includes multiple item samples, and each item sample includes item image data, text data for describing the item, and item style data, and where the item sample has a style label; For each item sample, input the item image data into the image feature extraction layer of the initial model to output item image features; input the text data for describing the item into the text feature extraction layer of the initial model to output text features; input the item style data into the style feature extraction layer of the initial model to output style features; Input the item image features and the text features into the feature combination layer of the initial model to output item combination features; Input the item combination features and the style features into the matching layer of the initial model to output a matching result for characterizing the matching between the item combination features and the style features; Adjust the model parameters of the initial model according to the matching result and the style label, input the multiple matching results and style labels output after adjusting the model parameters multiple times into the loss function, and when the change of the loss function approaches zero, obtain a trained feature recognition model; the loss function indicates the difference between the actual matching relationship between an item and a style and the matching relationship between the item and the style predicted by the algorithm.
2. The training method according to claim 1, wherein, Input the item image features and the text features into the image combination layer of the initial model, and the output of the item combination features includes: Use the image combination layer to splice the item image features and the text features into the item combination features.
3. The training method according to claim 1, wherein, After obtaining the item sample data set, it further includes: Generate an augmented item sample data set according to the item sample data set.
4. The training method according to claim 3, wherein, The generating the augmented item sample data set according to the item sample data set includes: According to the item image data, determine a first list of item data similar to the item image data from an item database, where the first list of item data includes item image data, text data for describing items, and item style data of different items; According to the text data for describing the item, determine a second list of item data similar to the text data for describing the item from the item database, where the second list of item data includes item image data, text data for describing items, and item style data of different items; Generate the augmented item sample data set according to the first list of item data and the second list of item data.
5. The training method according to claim 4, wherein, The generating the augmented item sample data set according to the first list of item data and the second list of item data includes: Generate the augmented item sample data set according to the data intersection of the first list of item data and the second list of item data.
6. The training method according to claim 1, wherein, The model parameters of the initial model include the model parameters of the image feature extraction layer of the initial model, the model parameters of the text feature extraction layer of the initial model, and the model parameters of the style feature extraction layer of the initial model. Adjusting the model parameters of the initial model according to the matching result and the style label to obtain the trained feature recognition model includes: Adjusting the model parameters of the image feature extraction layer of the initial model, the model parameters of the text feature extraction layer of the initial model, and the model parameters of the style feature extraction layer of the initial model according to the error between the matching result and the style label to obtain the trained feature recognition model.
7. An article feature recognition method, comprising: Obtaining the item information of the item to be processed, where the item information includes item image information and text information for describing the item; Inputting the item image information into the image feature extraction layer of the feature recognition model to output the item image features of the item to be processed, where the feature recognition model is trained by the training method described in any one of claims 1 to 6; Inputting the text information for describing the item into the text feature extraction layer of the feature recognition model to output the text features of the item to be processed; Inputting the item image features and the text features into the feature combination layer of the feature recognition model to output the item combination features of the item to be processed; Inputting the item combination features and the description features of the candidate style into the matching layer of the feature recognition model to output a matching result for characterizing the matching between the item combination features and the description features of the candidate style, where the description features of the candidate style are obtained after inputting the item style information retrieved from the candidate style database into the style feature extraction layer of the feature recognition model; Determining the style feature information that matches the item combination features of the item to be processed according to the matching result.
8. The method according to claim 7, further comprising: Obtaining a plurality of item review text data and a plurality of item title text data; Preprocessing the plurality of item review text data and the plurality of item title text data to obtain a text data set for characterizing the item style; Generating the candidate style database according to the text data set for characterizing the item style.
9. The method according to claim 8, further comprising: Generating a vector data set for characterizing the item style according to the text data set for characterizing the item style; Generating the candidate style database according to the vector data set for characterizing the item style.
10. A feature recognition model training apparatus, comprising: A first acquisition module for acquiring an item sample data set, where the item sample data set includes a plurality of item samples, and each item sample includes item image data, text data for describing the item, and item style data, and the item sample has a style label; A feature extraction module for, for each item sample, inputting the item image data into the image feature extraction layer of the initial model to output item image features; inputting the text data for describing the item into the text feature extraction layer of the initial model to output text features; and inputting the item style data into the style feature extraction layer of the initial model to output style features; A feature combination module, configured to input the item image features and the text features into a feature combination layer of the initial model, and output item combination features; A matching module, configured to input the item combination features and the style features into a matching layer of the initial model, and output a matching result for characterizing the matching between the item combination features and the style features; An adjustment module, configured to adjust model parameters of the initial model according to the matching result and the style label, input multiple matching results and style labels output after adjusting the model parameters multiple times into a loss function, and obtain a trained feature recognition model when the change of the loss function approaches zero; the loss function indicates the difference between the actual matching relationship between an item and a style and the matching relationship between the item and the style predicted by the algorithm.
11. An article feature recognition device, comprising: A second acquisition module, configured to acquire item information of an item to be processed, where the item information includes item image information and text information for describing the item; A first recognition module, configured to input the item image information into an image feature extraction layer of a feature recognition model, and output item image features of the item to be processed, where the feature recognition model is trained by the training method according to any one of claims 1 to 6; A second recognition module, configured to input the text information for describing the item into a text feature extraction layer of the feature recognition model, and output text features of the item to be processed; A third recognition module, configured to input the item image features and the text features into a feature combination layer of the feature recognition model, and output item combination features of the item to be processed; A fourth recognition module, configured to input the item combination features and the description features of a candidate style into a matching layer of the feature recognition model, and output a matching result for characterizing the matching between the item combination features and the description features of the candidate style, where the description features of the candidate style are obtained by inputting item style information acquired from a candidate style database into a style feature extraction layer of the feature recognition model; A determination module, configured to determine style feature information that matches the item combination features of the item to be processed according to the matching result; 12. An electronic device, comprising: One or more processors; A storage device, configured to store one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the method according to any one of claims 1 to 6 or any one of claims 7 to 9.
13. A computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor is caused to execute the method according to any one of claims 1 to 6 or any one of claims 7 to 9.
14. A computer program product, comprising a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 or any one of claims 7 to 9 is implemented.
Citation Information
Patent Citations
Copywriting generation method and device, copywriting evaluation model training method and device, and equipment
CN112232067A
Training method of video tag recommendation model and method for determining video tag
CN113378784A