Feature data generation method and device, article recommendation method and electronic equipment

By extracting the image features, text vector features and attribute features of automobile consumable parts products, and generating multi-feature data, the problem of low recognition accuracy of similar products in e-commerce platforms is solved and the order return rate is reduced.

CN120032136APending Publication Date: 2025-05-23BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311555711.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-21
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In the e-commerce sales scenario of automobile consumable parts, it is difficult for the existing technology to accurately identify the same product, resulting in a high order return rate.

Method used

By obtaining the item display image and page description text data, image feature extraction, text vector transformation and attribute entity recognition are performed, and multi-feature data of the target item is generated by combining selective search algorithms and pre-training models.

Benefits of technology

It improves the accuracy of identification of the same product and reduces the probability of order return and exchange. It is especially suitable for products with strong adaptability characteristics such as automobile consumable parts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032136A_ABST
    Figure CN120032136A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a feature data generation method and device, an article recommendation method and electronic equipment. According to one specific embodiment, the method comprises the steps that related data representing a target article on a platform are obtained, and the related data comprise an article display image and page description text data; performing feature extraction according to the article display image to obtain image features representing the target article; performing vector conversion according to the page description text data to obtain text vector features of the target article; performing attribute entity identification according to the page description text data, and obtaining attribute features of the target article based on the identified attribute data; and carrying out feature fusion on the image features, the text vector features and the attribute features to generate multi-feature data of the target object. The embodiment is related to a platform optimization technology, and the multi-dimensional feature data of the article can be extracted, so that the same article identification accuracy and the article recommendation accuracy of the article can be improved, and the probability of order refunding and changing can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of platform optimization technology, and specifically to a method and device for generating feature data, an item recommendation method, and an electronic device. Background Art

[0002] Automobiles are highly complex products. In the sales scenario of automobile consumable parts (parts that are easily damaged during the life cycle of the automobile or require regular replacement and maintenance), users or merchants hope to find a set of products with the same circulation attributes and usage functions as a target product from a massive pool of e-commerce products, that is, to find the same products as the target automobile consumable parts product. Therefore, how to improve the accuracy of identifying the same products of this type of products is an issue that needs to be solved urgently.

[0003] The above information disclosed in this Background section is only for enhancement of understanding of the background of the inventive concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the invention

[0004] The content of this disclosure is used to introduce concepts in a brief form, which will be described in detail in the detailed implementation section below. The content of this disclosure is not intended to identify the key features or essential features of the technical solution claimed for protection, nor is it intended to limit the scope of the technical solution claimed for protection.

[0005] Some embodiments of the present disclosure propose a feature data generation method, a feature data generation device, an item recommendation method, an item recommendation device, an electronic device, a computer-readable medium and a computer program product to solve one or more of the technical problems mentioned in the above background technology section.

[0006] In a first aspect, some embodiments of the present disclosure provide a method for generating feature data, including: acquiring relevant data representing a target item on a platform, wherein the relevant data includes an item display image and page description text data; performing feature extraction based on the item display image to obtain image features representing the target item; performing vector conversion based on the page description text data to obtain text vector features of the target item; performing attribute entity recognition based on the page description text data, and obtaining attribute features of the target item based on the recognized attribute data; and performing feature fusion on image features, text vector features, and attribute features to generate multi-feature data of the target item.

[0007] In some embodiments, feature extraction is performed based on an item display image to obtain image features representing the target item, including: using a selective search algorithm to generate a preset number of candidate areas on the item display image; obtaining the boundary dimensions of the preset number of candidate areas, normalizing the dimensions of each candidate area, and obtaining normalized candidate areas; classifying each normalized candidate area to determine a target candidate area representing the target item; correcting and adjusting the boundaries of the target candidate area, and extracting image features of the corrected target candidate area as image features of the target item.

[0008] In some embodiments, a selective search algorithm is used to generate a preset number of candidate regions on the object display image, including: generating multiple initial image regions on the object display image; merging adjacent initial image regions according to region similarity and a first threshold until all initial image regions are generated to obtain multiple initial segmented regions; and fusing multiple initial segmented regions according to the similarity between the initial segmented regions until a preset number of candidate regions are obtained.

[0009] In some embodiments, attribute entity recognition is performed based on page description text data, and attribute characteristics of the target item are obtained based on the recognized attribute data, including: inputting the page description text data into a pre-trained attribute extraction model to extract text attribute entities, wherein sample annotated attribute entities in sample data used to train the attribute extraction model are obtained by labeling attribute parameters in a standard item library; based on the standard item library, the text attribute entities are standardized and cleaned to obtain the attribute characteristics of the target item; and vector conversion is performed based on the page description text data to obtain text vector characteristics of the target item, including: using the word vector obtained after converting the text attribute entity vector as the text vector characteristic of the target item.

[0010] In some embodiments, standardized cleaning is performed on text attribute entities based on a standard item library, including: based on basic attributes of the target item, screening out a standard sub-item library that matches the basic attributes from the standard item library, wherein the basic attributes include brand logos and / or item categories; and standardized cleaning is performed on text attribute entities based on the standard sub-item library.

[0011] In some embodiments, the text attribute entity is standardized and cleaned according to the standard sub-item library, including: for each attribute in the text attribute entity, determining the similarity between the extracted attribute value and the corresponding standard attribute values ​​in the standard sub-item library; and determining whether to replace the extracted attribute value with the corresponding standard attribute value according to the attribute similarity and a second threshold.

[0012] In some embodiments, the text attribute entities are standardized and cleaned according to the standard item library, and the following steps are also included: the cleaned text attribute entities are consistency checked with the attribute parameters of the target item in the standard item library so that both attribute data represent the same item; and the verified consistent text attribute entities are used as the attribute features of the target item.

[0013] In some embodiments, feature fusion is performed on image features, text vector features, and attribute features to generate multi-feature data of the target object, including: converting image features into one-dimensional image features, and splicing the one-dimensional image features, text vector features, and attribute features to obtain multi-feature data of the target object.

[0014] In a second aspect, some embodiments of the present disclosure provide a feature data generating device, comprising: an acquisition unit, configured to acquire relevant data representing a target item on a platform, wherein the relevant data includes an item display image and a page description text data; an image feature extraction unit, configured to perform feature extraction based on the item display image to obtain image features representing the target item; a text feature conversion unit, configured to perform vector conversion based on the page description text data to obtain text vector features of the target item; an attribute feature recognition unit, configured to perform attribute entity recognition based on the page description text data, and obtain attribute features of the target item based on the recognized attribute data; and a feature fusion unit, configured to perform feature fusion on image features, text vector features and attribute features to generate multi-feature data of the target item.

[0015] In some embodiments, the image feature extraction unit is further configured to use a selective search algorithm to generate a preset number of candidate areas on the object display image; obtain the boundary dimensions of the preset number of candidate areas, normalize the dimensions of each candidate area, and obtain a normalized candidate area; classify each normalized candidate area to determine a target candidate area representing the target object; correct and adjust the boundaries of the target candidate area, and extract image features of the corrected target candidate area as image features of the target object.

[0016] In some embodiments, the image feature extraction unit is further configured to generate multiple initial image areas on the object display image; merge adjacent initial image areas according to area similarity and a first threshold until all initial image areas are generated to obtain multiple initial segmented areas; and fuse the multiple initial segmented areas according to the similarity between the initial segmented areas until a preset number of candidate areas are obtained.

[0017] In some embodiments, the attribute feature recognition unit is further configured to input the page description text data into a pre-trained attribute extraction model to extract text attribute entities, wherein the sample annotated attribute entities in the sample data used to train the attribute extraction model are obtained by labeling according to the attribute parameters in the standard item library; the text attribute entities are standardized and cleaned according to the standard item library to obtain the attribute features of the target item; and the text feature conversion unit is further configured to use the word vector obtained after converting the text attribute entity vector as the text vector feature of the target item.

[0018] In some embodiments, the attribute feature recognition unit is further configured to filter out a standard sub-item library that matches the basic attributes from the standard item library based on the basic attributes of the target item, wherein the basic attributes include brand logos and / or item categories; and perform standardized cleaning on the text attribute entities based on the standard sub-item library.

[0019] In some embodiments, the attribute feature recognition unit is further configured to determine, for each attribute in the text attribute entity, the similarity between the extracted attribute value and the corresponding standard attribute values ​​in the standard sub-item library; and determine whether to replace the extracted attribute value with the corresponding standard attribute value based on the attribute similarity and the second threshold.

[0020] In some embodiments, the attribute feature recognition unit is further configured to perform consistency verification on the cleaned text attribute entity and the attribute parameters of the target item in the standard item library so that both attribute data represent the same item; and use the verified consistent text attribute entity as the attribute feature of the target item.

[0021] In some embodiments, the feature fusion unit is further configured to convert the image features into one-dimensional image features, and to concatenate the one-dimensional image features, text vector features, and attribute features to obtain multi-feature data of the target object.

[0022] In a third aspect, some embodiments of the present disclosure provide an item recommendation method, comprising: in response to detecting an item search request, obtaining, from an item feature library, multi-feature data of a first item indicated by the item search request, wherein the multi-feature data in the item feature library is obtained using the feature data generation method described in any one of the implementations of the first aspect; determining the similarity between the multi-feature data of the first item and the multi-feature data of other items in the item feature library, and determining other items whose similarity is greater than a third threshold as candidate items for the first item; and generating item recommendation information based on the information of the first item and the candidate items for display.

[0023] In some embodiments, item recommendation information is generated based on information about a first item and candidate items, including: performing a consistency check on the attribute characteristics of the candidate items based on the attribute characteristics of the first item to determine whether the attribute characteristics of the candidate items are within a change range of the attribute characteristics of the first item; determining candidate items that are consistent with the check as items of the same type as the first item, and using the information of the first item and items of the same type as item recommendation information.

[0024] In a fourth aspect, some embodiments of the present disclosure provide an item recommendation device, comprising: a data acquisition unit, configured to, in response to detecting an item search request, acquire, from an item feature library, multi-feature data of a first item indicated by the item search request, wherein the multi-feature data in the item feature library is obtained using the feature data generation method described in any one of the implementations of the first aspect above; a feature similarity determination unit, configured to determine the similarity between the multi-feature data of the first item and the multi-feature data of other items in the item feature library, and to determine other items whose similarity is greater than a third threshold as candidate items for the first item; and a recommendation information generation unit, configured to generate item recommendation information based on the information of the first item and the candidate items for display.

[0025] In some embodiments, the recommendation information generation unit is further configured to perform consistency verification on the attribute characteristics of the candidate item based on the attribute characteristics of the first item to determine whether the attribute characteristics of the candidate item are within the change range of the attribute characteristics of the first item; determine the candidate items that are consistent with the verification as the same type of items as the first item, and use the information of the first item and the same type of items as item recommendation information.

[0026] In a fifth aspect, some embodiments of the present disclosure provide an electronic device, comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect or the third aspect above.

[0027] In a sixth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the method described in any one of the implementations in the first aspect or the third aspect is implemented.

[0028] In a seventh aspect, some embodiments of the present disclosure provide a computer program product, including a computer program, which, when executed by a processor, implements the method described in any implementation of the first aspect or the third aspect.

[0029] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: the feature data generation method of some embodiments of the present disclosure can generate multi-dimensional feature data of items, which helps to improve the accuracy of same-product identification and thus reduce the probability of order returns. Specifically, in the related same-product identification method, the feature data of the goods is mainly obtained by using the image extraction and text extraction conversion in the general field, so as to match the goods to realize the identification and judgment of similar goods. However, the application field of this technology is too generalized, and the main scope is generally the general commodity domain. The category characteristics and adaptation expertise of the goods are not taken into account, especially the categories with adaptation characteristics and compatibility characteristics behind the purchase of goods, such as the specification attributes and extended attribute content of the goods in the field of automotive wearing parts. In other words, this method does not make full use of the multi-dimensional and multi-modal product feature information of the goods. This will affect the accuracy of the same-product identification of such goods, thereby increasing the number of return and exchange orders.

[0030] Based on this, the feature data generation method of some embodiments of the present disclosure can extract the product main image features, product text description features, and category-specific specification attribute features of the product to obtain multi-feature data of the product to identify the same product. Thereby, the accuracy of the product identification results can be guaranteed and the accuracy of identification can be improved. In particular, in the scenario of automobile consumable parts sold on e-commerce platforms, the multi-dimensional and multi-modal product feature information of the product can be fully utilized to meet the "product-model" strong adaptation scenario requirements. This helps to reduce the probability of order returns and exchanges. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.

[0032] Figure 1 are flow charts of some embodiments of the characteristic data generating method disclosed herein;

[0033] Figure 2A is a flow chart of some embodiments of image feature extraction;

[0034] Figure 2B are flow charts of other embodiments of image feature extraction;

[0035] Figure 3A is a flow chart of some embodiments of text feature and attribute feature extraction;

[0036] Figure 3B is a flow chart of some embodiments of attribute standardization cleaning;

[0037] Figure 3C is a flow chart of some embodiments of attribute consistency checking;

[0038] Figure 4 It is a schematic diagram of the structure of some embodiments of the characteristic data generating device disclosed in the present invention;

[0039] Figure 5 is a flow chart of some embodiments of the item recommendation method disclosed in the present invention;

[0040] Figure 6 is a schematic diagram of the structure of some embodiments of the item recommendation device disclosed in the present invention;

[0041] Fig. 7A are schematic diagrams of some application scenarios applicable to the disclosed method;

[0042] Figure 7B It is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0043] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0044] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.

[0045] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0046] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0047] Figure 1 The process 100 of some embodiments of the feature data generation method according to the present disclosure is shown. The method may include the following steps:

[0048] Step 101, obtaining relevant data representing a target object on a platform.

[0049] In some embodiments, the execution subject (such as a data processing server) of the feature data generation method can be connected to other devices (such as a platform server) by wired connection or wireless connection. Here, the execution subject can obtain relevant data representing the target item on the platform from the platform database. Among them, the relevant data may include item display images and page description text data. The item display image can be a picture or video used to display the appearance (such as shape, color), structural structure, etc. of the item, such as the main picture of the item in the product details page. The page description text data can be text data describing the detailed information of the item on the page, such as text description data in the product details page. The target items here can be various items displayed on the platform, especially items with strong adaptability requirements to the use or application scenarios, such as automotive consumables, home appliances or electronic product accessories.

[0050] Step 102: extract features based on the object display image to obtain image features representing the target object.

[0051] In some embodiments, the execution subject may perform feature extraction based on the acquired item display image to obtain image features that characterize the target item. Here, the feature extraction method may be selected based on actual conditions, such as inputting the item display image into a CNN (convolutional neural network) to predict the category and location of different items, and then extracting image features of the location of the target item.

[0052] It should be noted that pictures are often an important carrier for obtaining product information, and pictures have different angles, sizes, colors, etc. when they are taken. Therefore, it is necessary to use a unified scale to represent the characteristics of pictures for comparison and application of product main pictures.

[0053] Optionally, the execution subject may use a selective search algorithm to generate a preset number of candidate regions on the object display image. Then, the boundary sizes of the preset number of candidate regions are obtained, and the sizes of each candidate region are normalized to obtain a normalized candidate region. After that, each normalized candidate region is classified to determine a target candidate region representing the target object. Finally, the boundary of the target candidate region is corrected and adjusted, and the image features of the corrected target candidate region are extracted as the image features of the target object.

[0054] For details, please refer to Figure 2A and Figure 2BThe processing flow is shown in the figure. Here, taking automobile wearing parts as an example, the main purpose of target detection is to distinguish the content that is most likely to be the main body of the product from the picture of wearing parts. The basic idea is to use the candidate box generation algorithm to select several pending grids. That is, use multiple rectangular boxes to circle all possible image objects. Because there may be occlusions between objects, the selected objects may be repeated. After feature extraction of the candidate box, it is sent to the bounding box regression network for classification. Then detection is performed to obtain the detection result. The process of two-stage target detection can be as follows: Figure 2A shown.

[0055] In the two stages of this application, the target detection model R-CNN can be used to detect the target of the main product image. It should be noted that the traditional method generally generates the image candidate frame by exhaustively enumerating the input image. This often generates a lot of useless candidate frames, which greatly wastes computing resources and computing efficiency. Therefore, in this embodiment, the optimized R-CNN can be used to perform target detection.

[0056] Here, R-CNN can use the selective search algorithm to perform efficient image candidate frame screening through image segmentation methods. As an example, multiple initial image regions are first generated on the item display image. Then, the adjacent initial image regions can be merged according to the region similarity and the first threshold until all the initial image regions are generated to obtain multiple initial segmented regions. That is, all the initial image regions are traversed. Afterwards, multiple initial segmented regions are fused according to the similarity between the initial segmented regions until a preset number of candidate regions are obtained.

[0057] Specifically, the first step is to use a greedy method to group and iterate the image regions. In this process, the differences caused by the scene, lighting, etc. are also filtered, which improves the speed and efficiency of selecting candidate boxes:

[0058] a) Generate an initial image region using, for example, Felzenszwalb and Huttenlocher (an image segmentation algorithm);

[0059] b) Calculate the similarity between the target and the initial area;

[0060] c) combining regions whose region similarity is greater than a threshold;

[0061] d) further calculating the similarity between the merged region and the adjacent regions, and merging the data;

[0062] e) Repeat the process of c and d until the image becomes a whole area, in which all the initial segmented areas are obtained;

[0063] f) Initialize and obtain all region sets R set = {r 1 ,r 2 ,…,r n};

[0064] g) Iterate through the region set R set ,calculate times, calculate the r selected each time i ,r j Similarity: merge similarity regions and iterate until the set number of candidate regions L is obtained;

[0065] h) Finally, the boundary results of the L candidate regions are output.

[0066] Step 2: Next, for the extracted candidate regions, the network parameters trained on the ImageNet (visual database) network can be used as extractors to extract data features of all bounding box result images. The extracted feature sizes are normalized and used as feature input for subsequent classification and bounding box regression training.

[0067] Step 3: Finally, the border is fine-tuned and corrected through the regression model to make the border positioning more accurate, ensure that the useless background is removed more accurately, and obtain accurate target detection results. In this example, the main body of the main image of the automobile wearing parts product is obtained.

[0068] Furthermore, if Figure 2B As shown in the figure, after the product main image target detection model is used, the main content of the main image of the automobile wearing parts product can be obtained. In order to ensure the rich extraction of image features and the constant output dimension, the image needs to be vectorized and converted into image recall features.

[0069] In some embodiments, the VGG16 model can be selected to remove the last three layers of fully connected networks, that is, the first 13 layers of convolutional networks, as the extraction model for feature conversion to obtain the original extracted features Vori (ie, V_ori). At this time, the extracted features can be used as the image features of the target object.

[0070] Optionally, in order to facilitate subsequent feature fusion and multi-feature recall with features other than image features, you can use compressed tensor methods such as flatten or reshape to straighten the extracted original features and convert them into one-dimensional image features. You can use V p (i.e. V_p) to represent it, waiting to be fused with subsequent text data and recalled.

[0071] Step 103, performing vector conversion based on the page description text data to obtain the text vector features of the target item.

[0072] In some embodiments, the execution subject may perform vector conversion based on the page description text data to obtain the text vector features of the target item. As an example, the execution subject may use a text-to-text vector conversion model, input the page description text data into the model, and output a corresponding text vector as the text vector feature of the target item.

[0073] Optionally, the execution subject may also use a pre-trained item text conversion and attribute extraction model to obtain the text vector features and attribute features of the target item. For details, please refer to the relevant description of step 104. Among them, the item text conversion and attribute extraction model mainly completes the extraction of the full text feature vector V from the product description text. t , and extract the information of automobile commodity attributes (extended attributes, specification attributes, etc.) of wearing parts contained in the text.

[0074] Step 104 , performing attribute entity recognition according to the page description text data, and obtaining attribute features of the target item based on the recognized attribute data.

[0075] In some embodiments, the execution subject may use various entity recognition methods to perform attribute entity recognition on the page description text data. Afterwards, the attribute characteristics of the target item may be obtained based on the recognized attribute data. For example, the recognized attribute data (i.e., attribute entity words) may be used as attribute characteristics. The attribute entity here is usually an entity word that characterizes the attributes of the item. The attribute characteristics may be feature data that characterizes the attributes of the item, such as specification attributes, extended attributes (such as adaptation characteristics, compatibility characteristics, etc.), etc.

[0076] Optionally, in order to improve the accuracy of the recognition result, the execution entity can input the page description text data into a pre-trained attribute extraction model, namely, an item text conversion and attribute extraction model, to extract text attribute entities. Among them, the sample annotation attribute entities in the sample data used to train the attribute extraction model are obtained by labeling according to the attribute parameters in the standard item library.

[0077] As an example, Figure 3A As shown, the attribute extraction model can take the BERT (Bidirectional Encoder Representations from Transformers, language representation model) model as an example, which will involve data labeling, model fine-tuning training, and model prediction to obtain text attribute entities and text word vectors.

[0078] The first step is to prepare for the text annotation of automobile wearing parts data. The BIO annotation method can be used to pre-annotate the attribute entities in the data text. The annotated attribute entities are the entity words that need to be recognized by the model. For example, for automobile wipers, the attribute entities that need to be annotated may include "boneless", "double-layer", "rubber strip", "silent", "steel sheet", "universal type", etc. Here, you can use the regular matching method and the attribute information of the standard commodity library to pre-BIO annotate the text. During the pre-annotation process, there may be missing, erroneous, and repeated annotation results. Afterwards, the data can be manually verified and corrected for a second time, such as by data operation personnel in the automobile wearing parts category. This completes the text annotation and preparation of the BIO data.

[0079] The next step is to divide the labeled data into data sets. All data can be divided into training data and test data in a ratio of 8:2 (which can be set by yourself). And it is processed into the corresponding sentence-label pair format for the downstream commodity attribute entity word training of the BERT pre-training model. And some parameters for model training can be set, such as the number of full sample training (num_train_epochs), batch training data volume (train_batch_size), maximum text length (max_seq_length), learning rate (learning_rate), etc. During the training process, the F1 value (an indicator of the accuracy of the binary classification model), accuracy (precision), recall rate (recall) and other indicators can be used to evaluate the model entity word extraction effect. Through the grid adjustment of the hyperparameters, the model with the best evaluation effect is selected as the final model for saving.

[0080] After the model is saved, the final attribute entity extraction result of the model can be used as the text attribute entity E = {e 1 ,e 2 ,…,e n}. At the same time, the word vector output by the last layer of BERT can be selected as the word vector V of the whole text t That is, the word vector obtained by converting the text attribute entity vector is used as the text vector feature of the target item.

[0081] In some embodiments, the attribute data (extended attributes and specification attributes) maintained by the items may be maintained in a non-standard manner or with deviations. This will lead to deviations and inaccuracies in the identification process of the same product. Therefore, in order to improve the accuracy of the recognition results, the text attribute entities can be standardized and cleaned according to the standard item library to obtain the attribute characteristics of the target item. That is, the reasonable extended attributes and specification attributes in the standard library can be used for normalization, and the non-standard attributes can be cleaned. As an example, the attribute parameters of the target item in the standard item library can be used to replace the extracted attribute values.

[0082] In some application scenarios, in order to make full use of the data in the standard item library, the attribute parameters of items of the same category and / or the same brand can also be used for standardized cleaning. Specifically, first, based on the basic attributes of the target item, a standard sub-item library that matches the basic attributes can be screened out from the standard item library. Among them, the basic attributes may include brand logos and / or item categories. Then, the text attribute entity can be standardized and cleaned based on the standard sub-item library. That is, for each attribute in the text attribute entity, the similarity between the extracted attribute value and the corresponding standard attribute values ​​in the standard sub-item library can be determined. Furthermore, based on the attribute similarity and the second threshold, it is determined whether to replace the extracted attribute value with the corresponding standard attribute value.

[0083] As an example, obtain the extended attributes and specification attribute set of all product maintenance as E x , which contains n attributes:

[0084]

[0085] For a product i, the attribute set is E i =[e i1 ,e i2 ,…,e im ,…,e in ]. Among them, some product attributes may be missing, and the missing items can be left blank. In addition to extended attributes and specification attributes, two other standardized basic attributes are obtained from the basic attributes of the product: brand information and product category, named Attr b and Attr c .

[0086] For product i, obtain its brand and product category, and the extended attributes and specification attribute set E maintained in all products x In the example above, use Attr b and Attr c , filter out the extended attributes and specification attribute subsets that meet the brand and product category Use E againi Each specific attribute value in is named e im , the extended attributes and specification attribute subset E that meet the brand and product category i x Each item in e jk Calculate the similarity of attributes:

[0087]

[0088] Among them, Levenshtein is the calculation method of edit distance, and the extended attributes and specification attribute subsets that meet the brand and product category are obtained. List of all Levenshtein similarity edit distance scores in . See Figure 3B As shown, a similarity threshold can be set, and the top N attributes with similarity greater than the threshold can be used as converted attributes to update and replace the original non-standardized attribute values ​​for consistency verification and feature encoding in subsequent processes. This can fully consider the situation of inconsistent equivalent meanings and inconsistent attribute descriptions, and standardize the attributes. If the similarity is not greater than the threshold, the original attribute value can be maintained. It should be noted that the same attribute may match multiple attribute values ​​that meet the threshold, so that users can compare and select.

[0089] Furthermore, in order to ensure the consistency of the items, the execution subject can also perform consistency verification of the attributes. That is, the cleaned text attribute entity can be checked for consistency with the attribute parameters of the target item in the standard item library, so that the two attribute data represent the same item. Then, the verified consistent text attribute entity can be used as the attribute feature of the target item.

[0090] Here, if Figure 3C As shown, the verification method can be performed using a parallel multi-attribute simultaneous verification method. There are two parts of attribute data in total: a) the product description attribute result extracted from the text, and b) the attribute result of the product maintenance. N groups (A1-An) of data to be verified are obtained in parallel, and consistency verification is performed to ensure that the results of the two parts of data to be verified are consistent.

[0091] As an example, the product attributes extracted from the description text of automobile consumable parts and the maintenance attributes of automobile consumable parts need to be consistent, because the attributes of the two dimensions refer to the same product. When the product attributes extracted from the description text are inconsistent with the attributes of the product itself, there will be inconsistencies with the target model car when used, thereby increasing customer complaints or return rates. Therefore, when the attributes are inconsistent, the merchant product governance plan can be initiated to ensure the consistency of the product attributes by modifying the maintenance text or maintenance attributes. For example, one attribute of a product is A1 B1, and the other attribute is A2. At this time, the corresponding product is A, and the B1 attribute value can be removed.

[0092] It is understandable that since the attributes of automobile wearing parts are relatively few and relatively standardized, all features can be encoded using sequence encoding. That is, a mapping relationship table can be maintained separately, and the attribute features Attr all ={A 1 ,A 2 ,…,A n} is encoded from 1 to n. Among them, 0 can be used to reserve the attribute value that is not maintained by the encoding. At this time, the mapping relationship table can also be used to feature encode the text attribute entities that have been verified to be consistent, so as to serve as the attribute characteristics of the target item. That is, the longest attribute combination length can be used as the maximum length of the attribute vector, and the feature encoding can be added in the order of appearance to obtain the attribute feature vector V a This can reduce the amount of attribute feature data, thereby reducing storage space usage, and also help improve the processing efficiency of subsequent item matching.

[0093] It should be noted that if the extracted text attribute entity adopts the feature encoding method, the attribute value maintained by the product can also be mapped in a normative manner during the consistency check. Compared with text verification, encoding verification is simpler and is conducive to improving the verification comparison efficiency.

[0094] Step 105, performing feature fusion on the image features, text vector features and attribute features to generate multi-feature data of the target object.

[0095] In some embodiments, the execution subject may perform feature fusion on the image features, text vector features and attribute features obtained in the above steps to generate multi-feature data of the target object, such as fusion in a certain feature order.

[0096] It should be noted that image features are usually multi-dimensional (such as three-dimensional) data. Therefore, in order to be able to merge them with text vector features and attribute features, the image features can be first converted into one-dimensional image features. Then the one-dimensional image features, text vector features and attribute features are spliced ​​to obtain multi-feature data of the target item. For example, the multi-feature vector data V of a certain product can be obtained by splicing them in the current order. all :

[0097] V all =concat(flatten(V p )+flatten(V t )+flatten(V a ));

[0098] In some embodiments, the method in the above embodiment can be used to initialize and store multi-feature data for all items, thereby obtaining an item feature library. When a new item appears on the platform, the multi-feature data of the item can be calculated and updated to the item feature library.

[0099] Through the above description, the feature data generation method of some embodiments of the present disclosure can extract the main image features of the product, the text description features of the product, and the category-specific specification attribute features of the product to obtain multi-feature data of the product, so as to identify the same product of the product. Thereby, the accuracy of the identification results of the same product can be guaranteed and the accuracy of identification can be improved. In particular, in the scenario of automobile consumable parts sold on e-commerce platforms, the multi-dimensional and multi-modal product feature information of the product can be fully utilized to meet the "product-model" strong adaptation scenario requirements. This helps to reduce the probability of order returns and exchanges.

[0100] Further references Figure 4 , as a response to the above Figure 1 In order to realize the method shown in the figure, the present disclosure provides some embodiments of a feature data generating device. These device embodiments are similar to Figure 1 The device can be specifically applied to various electronic devices.

[0101] like Figure 4As shown, the feature data generating device 400 of some embodiments may include: an acquisition unit 401, configured to acquire relevant data representing the target item on the platform, wherein the relevant data includes item display images and page description text data; an image feature extraction unit 402, configured to perform feature extraction based on the item display images to obtain image features representing the target item; a text feature conversion unit 403, configured to perform vector conversion based on the page description text data to obtain text vector features of the target item; an attribute feature recognition unit 404, configured to perform attribute entity recognition based on the page description text data, and obtain attribute features of the target item based on the recognized attribute data; a feature fusion unit 405, configured to perform feature fusion on image features, text vector features and attribute features to generate multi-feature data of the target item.

[0102] In some embodiments, the image feature extraction unit 402 can be further configured to use a selective search algorithm to generate a preset number of candidate areas on the object display image; obtain the boundary dimensions of the preset number of candidate areas, normalize the dimensions of each candidate area, and obtain a normalized candidate area; classify each normalized candidate area to determine a target candidate area representing the target object; correct and adjust the boundaries of the target candidate area, and extract the image features of the corrected target candidate area as the image features of the target object.

[0103] In some embodiments, the image feature extraction unit 402 can be further configured to generate multiple initial image areas on the object display image; merge adjacent initial image areas according to the area similarity and the first threshold until all the initial image areas are generated to obtain multiple initial segmented areas; and fuse the multiple initial segmented areas according to the similarity between the initial segmented areas until a preset number of candidate areas are obtained.

[0104] In some embodiments, the attribute feature recognition unit 403 can be further configured to input the page description text data into a pre-trained attribute extraction model to extract text attribute entities, wherein the sample annotation attribute entities in the sample data used to train the attribute extraction model are obtained by labeling according to the attribute parameters in the standard item library; based on the standard item library, the text attribute entities are standardized and cleaned to obtain the attribute features of the target item; and the text feature conversion unit 404 can be further configured to use the word vector obtained after converting the text attribute entity vector as the text vector feature of the target item.

[0105] In some embodiments, the attribute feature recognition unit 403 can be further configured to filter out a standard sub-item library that matches the basic attributes from the standard item library based on the basic attributes of the target item, wherein the basic attributes include brand logos and / or item categories; and perform standardized cleaning on the text attribute entities based on the standard sub-item library.

[0106] In some embodiments, the attribute feature recognition unit 403 can be further configured to determine, for each attribute in the text attribute entity, the similarity between the extracted attribute value and the corresponding standard attribute values ​​in the standard sub-item library; and determine whether to replace the extracted attribute value with the corresponding standard attribute value based on the attribute similarity and the second threshold.

[0107] In some embodiments, the attribute feature recognition unit 403 can also be configured to perform consistency verification on the cleaned text attribute entity and the attribute parameters of the target item in the standard item library so that both attribute data represent the same item; and use the verified consistent text attribute entity as the attribute feature of the target item.

[0108] In some embodiments, the feature fusion unit 405 may be further configured to convert image features into one-dimensional image features, and to concatenate the one-dimensional image features, text vector features, and attribute features to obtain multi-feature data of the target object.

[0109] It is understandable that the units described in the feature data generating device 400 are similar to those described in the reference Figure 1 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the feature data generating device 400 and the units included therein, and will not be described in detail here.

[0110] In some application scenarios, based on Figure 1 The multi-feature data of each item obtained in the embodiment can be stored to obtain an item feature library. The item feature library can then be used to perform item matching queries, especially for those items with multi-dimensional and multi-modal commodity feature information, as well as strong adaptability and compatibility characteristics, such as automobile consumable parts.

[0111] like Fig. 7A In the application scenario shown, the technical solution of the embodiment of the present disclosure can solve the problem of identifying the same product of automobile wearing parts, making the identification accuracy more accurate. The implementation process can be divided into the following nine steps:

[0112] Step 1: The system user enters the product number of the product to be queried;

[0113] Step 2: The system loads the product main image, product detail page description text, product specification attribute extended attributes and other features of the target product to be matched from the product storage system, as well as the same feature data of the product pool of the target category to be matched;

[0114] Step 3: For the product main image data, use the object detection model and graphic feature extraction model trained based on the main image data of automobile wearing parts to extract image features for characterizing the product image features;

[0115] Step 4: For the target product’s product details page description text data, perform feature extraction on two parts. The first part is the full text feature, and the second part is the extended attributes and specification attributes:

[0116] a) The first part is to transform the features of the description text on the product details page, using the text-to-text vector conversion model trained based on the text of automobile wearing parts products to extract the text vector features of the full text dimension;

[0117] b) The second part is to extract the specification attributes and special attributes of the description text on the product details page. The specification attributes and special attributes extraction model of automobile consumable parts category trained based on automobile consumable parts data is used to extract the specification attributes and special attributes of the products in the text;

[0118] Step 5: Perform consistency check on the product itself. In step 4b), the named entity extraction model is used to extract product specification attributes and extended attribute information from the product details page text. To ensure the consistency of product specification attributes and extended attributes when the product is released, a consistency check is performed on the product maintenance attributes. The consistency between the specification attributes and extended attributes of the product when it is released and the text description of the product details page is ensured to pass the check.

[0119] Step 6: Convert the verified specification attributes and extended attributes into attribute features through feature coding;

[0120] Step 7: Through the first six steps, the product main image features, product details page text features, and consistency check attribute features can be obtained, and the features are integrated. At the same time, the processing flow of steps 2 to 7 is carried out for all product data, and all data that meets the conditions is converted into multi-feature data and stored in the feature storage engine, which is convenient for recall and use in subsequent steps.

[0121] Step 8: Use the multimodal features of the target product that are formed by integrating the product main image features, the product detail page text features, and the attribute features that meet the consistency check, and use the vector similarity calculation retrieval method to calculate the feature similarity with the multimodal features in the full pool of products in the same category to obtain a coarse-grained recall product result set;

[0122] Step 9: For the coarse-grained recalled product results, verify the consistency of the recalled products before output, and output the same product identification results for downstream applications.

[0123] For specific product matching recall procedures, please refer to Figure 5 The relevant description of the embodiments will not be repeated here.

[0124] See below Figure 5 , which shows a process 500 of some embodiments of the item recommendation method according to the present disclosure. The method comprises the following steps:

[0125] Step 501, in response to detecting an item search request, obtaining multi-feature data of a first item indicated by the item search request from an item feature library.

[0126] In some embodiments, the execution subject of the item recommendation method (such as a platform server) can communicate with other electronic devices through a wired connection or a wireless connection. Here, the user can perform various operations on the platform through the terminal device, such as entering an item keyword or number, etc., to search for the required item. At this time, when the execution subject detects an item search request, it can obtain the multi-feature data of the first item indicated by the item search request from the item feature library. Among them, the multi-feature data in the item feature library is usually obtained by the above Figure 1 The feature data generation method described in any implementation manner in the embodiment is obtained.

[0127] As an example, the item feature library may store the correspondence between item identifiers (such as names, codes, etc.) and multi-feature data. In this way, the execution subject may obtain the multi-feature data corresponding to the item (i.e., the first item) from the item feature library through the item name, keywords (such as attributes and uses, etc.) or item code in the item search request.

[0128] It should be noted that the execution subject in this embodiment can be Figure 1 The execution subjects in the embodiments may be the same subject or may be different subjects.

[0129] Step 502, determining the similarity between the multi-feature data of the first object and the multi-feature data of other objects in the object feature library, and determining the other objects with similarity greater than a third threshold as candidate objects of the first object.

[0130] In some embodiments, the execution subject may determine the similarity between the multi-feature data of the first item and the multi-feature data of other items in the item feature library. The other items are usually items other than the first item. That is, the similarity between the two multi-feature data is calculated. The similarity calculation method here can be set according to the actual situation. As an example, the angle cosine or the Euclidean distance deformation formula can be used to represent the multi-feature vectors V of the two commodities. x and V y Similarity score:

[0131]

[0132] Among them, v xi 、v yi Represent multiple eigenvectors V x 、V y The i-th feature in .

[0133] Furthermore, the execution subject may determine other items with a similarity greater than the third threshold as candidate items for the first item. The third threshold here can also be set by oneself. In other words, for the multi-feature data (vector) of each item in the item feature library, the degree of similarity they point to the physical object can be determined according to the degree of similarity. The vector recall measurement is performed through the similarity score, and by controlling the similarity threshold, a set of all coarse-grained commodity pools that meet the threshold condition is obtained.

[0134] It should be noted that as the number of automobile consumable parts increases, the vector recall engine can also be used to recall vectors of similar features of products. For example, the Faiss engine can be used to store features and recall the most relevant TopN vectors, and the following is obtained: Fig. 7A The coarse-grained product pool set shown in . Among them, Faiss (Facebook AI SimilaritySearch) is usually an efficient similarity search library.

[0135] Step 503: Generate item recommendation information based on the information of the first item and the candidate items for display.

[0136] In some embodiments, the execution entity may generate item recommendation information based on the information of the first item and the candidate items for display. This allows for downstream applications, such as user price comparison, merchant price adjustment for their own products, and the platform's introduction of a competition mechanism for products in the same category. As an example, the execution entity may directly use the information of the first item and the information of the candidate items as item recommendation information. Alternatively, the information of the first item and the information of several candidate items with the highest similarity may be used as item recommendation information.

[0137] In some embodiments, in order to improve the accuracy of the recommendation results, the execution subject may also perform an attribute consistency check. Specifically, the attribute characteristics of the candidate items may be checked for consistency based on the attribute characteristics of the first item to determine whether the attribute characteristics of the candidate items are within the attribute characteristic change range of the first item. The candidate items that are consistent with the check are determined to be the same product as the first item. At this time, the information of the first item and the same product item may be used as item recommendation information.

[0138] That is to say, for Fig. 7A In the application scenario shown, after obtaining the coarse-grained product matching set, the output results can be checked for consistency. Here, the main purpose is to check whether the specification attributes and extended attributes of automobile consumable parts meet the reasonable change range space of the specification attributes and extended attributes of the queried product, as well as the upper and lower compatibility relationship. This information can also be recorded in the standard item library in advance. Then, the coarse-grained recall results that do not meet the above conditions are discarded. The consumable parts that meet the consistency check results are used as the fine-grained product pool that meets the conditions, and are output as the final matching same-product recognition results of the model.

[0139] It is understandable that if the attribute features in the multi-feature data are coded data, the attribute features can be converted according to the previous mapping relationship table before the consistency check.

[0140] The item recommendation method of the disclosed embodiment mainly solves the application scenarios where the goods and usage scenarios have strong adaptability, especially the goods of automobile wearing parts. The applicable fields of the prior art are too generalized, and the category characteristics and adaptation expertise of the goods are not taken into account. In particular, the categories behind the purchase of goods have adaptation characteristics and compatibility characteristics. That is, the specification attributes and extended attribute content of the goods in the field of automobile wearing parts are not taken into account.

[0141] Therefore, in the field of automobile consumable parts product matching, the method disclosed in the present invention adds a specification attribute and extended attribute feature based on product maintenance for product recall on the basis of optimizing the two feature recalls of the original picture and full text. At the same time, the verification work of attribute consistency extraction, conversion and matching of the product itself based on named entity extraction is added. And the output results are checked for consistency with the same product. The above added innovations are mainly to meet the important "product model-car model" strong adaptation logic in the field of automobile consumable parts, ensuring that the same product found is more accurate. It meets the application scenarios under strict adaptation scenarios, making the found same products more accurate, thereby reducing the number and probability of order returns.

[0142] In addition, for users, the results of similar product identification can improve their shopping experience, compare product prices based on the identified similar products, and purchase products that meet their current budget. For merchants, the results of similar product identification can improve the competitiveness of their products, set reasonable competitive prices for their products, and increase sales and exposure. For platforms, the results of similar product identification can introduce a similar product competition mechanism to ensure price discounts within the platform, return profits to consumers, and improve conversions.

[0143] Further references Figure 6 , as a response to the above Figure 5 The present disclosure provides some embodiments of an item recommendation device. These device embodiments are similar to Figure 5 The device can be specifically applied to various electronic devices.

[0144] like Figure 6 As shown, the item recommendation device 600 of some embodiments may include: a data acquisition unit 601, configured to, in response to detecting an item search request, acquire multi-feature data of a first item indicated by the item search request from an item feature library, wherein the multi-feature data in the item feature library is obtained by using the above Figure 1 The feature data generation method described in any implementation manner in the embodiment is obtained; the feature similarity determination unit 602 is configured to determine the similarity between the multi-feature data of the first item and the multi-feature data of other items in the item feature library, and to determine other items with a similarity greater than a third threshold as candidate items for the first item; the recommendation information generation unit 603 is configured to generate item recommendation information based on the information of the first item and the candidate items for display.

[0145] In some embodiments, the recommendation information generation unit 603 can be further configured to perform consistency verification on the attribute characteristics of the candidate items based on the attribute characteristics of the first item to determine whether the attribute characteristics of the candidate items are within the change range of the attribute characteristics of the first item; determine the candidate items that are consistent with the verification as the same type of items as the first item, and use the information of the first item and the same type of items as item recommendation information.

[0146] It is understandable that the units recorded in the item recommendation device 600 are similar to those in the reference Figure 5 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the item recommendation device 600 and the units included therein, and will not be described in detail here.

[0147] Reference below Figure 7B , which shows a structural schematic diagram of an electronic device 700 suitable for implementing some embodiments of the present disclosure. Figure 7B The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0148] like Figure 7B As shown, the electronic device 700 may include a processing device 701 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the terminal device 700 are also stored. The processing device 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0149] Typically, the following devices may be connected to the I / O interface 705: input devices 706 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 707 including, for example, a speaker, a vibrator, etc.; storage devices 708 including, for example, a disk, a hard disk, etc.; and communication devices 709. The communication devices 709 may allow the electronic device 700 to communicate with other devices wirelessly or by wire to exchange data. Figure 7B The electronic device 700 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 7B Each block shown in the figure may represent one device, or may represent multiple devices as required.

[0150] In particular, according to some embodiments of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from the network through the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above-mentioned functions defined in the method of some embodiments of the present disclosure are executed.

[0151] It should be noted that the computer-readable medium recorded in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0152] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0153] The computer-readable medium may be included in the electronic device; or it may exist independently without being installed in the electronic device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: obtains relevant data representing the target item on the platform, wherein the relevant data includes item display images and page description text data; performs feature extraction based on the item display images to obtain image features representing the target item; performs vector conversion based on the page description text data to obtain text vector features of the target item; performs attribute entity recognition based on the page description text data, and obtains attribute features of the target item based on the recognized attribute data; performs feature fusion on image features, text vector features, and attribute features to generate multi-feature data of the target item.

[0154] Alternatively, the electronic device: in response to detecting an item search request, obtains from an item feature library the multi-feature data of the first item indicated by the item search request, wherein the multi-feature data in the item feature library is obtained by using the above Figure 1 The feature data generation method described in any implementation manner in the embodiments is obtained; determining the similarity between the multi-feature data of the first item and the multi-feature data of other items in the item feature library, and determining other items with a similarity greater than a third threshold as candidate items for the first item; generating item recommendation information based on the information of the first item and the candidate items for display.

[0155] In addition, computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0156] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0157] The units described in some embodiments of the present disclosure may be implemented by software or by hardware. The described units may also be provided in a processor, for example, may be described as: a processor including an acquisition unit, an image feature extraction unit, a text feature conversion unit, and an attribute feature recognition unit; or including a data acquisition unit, a feature similarity determination unit, and a recommendation information generation unit. The names of these units do not, in some cases, constitute limitations on the units themselves, for example, the acquisition unit may also be described as a "unit for acquiring relevant data characterizing a target item on a platform".

[0158] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0159] Some embodiments of the present disclosure further provide a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned feature data generation methods or item recommendation methods.

[0160] The above descriptions are only some preferred embodiments of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) and the technical solutions formed.

Claims

1. A method for generating feature data, include: Acquire relevant data representing the target item on the platform, wherein the relevant data includes item display images and page description text data; Extracting features from the displayed image of the object to obtain image features representing the target object; Performing vector conversion according to the page description text data to obtain text vector features of the target item; Performing attribute entity recognition according to the page description text data, and obtaining attribute characteristics of the target item based on the recognized attribute data; The image features, the text vector features and the attribute features are fused to generate multi-feature data of the target object.

2. The feature data generating method according to claim 1, in, The extracting features according to the object display image to obtain image features representing the target object includes: Using a selective search algorithm, a preset number of candidate regions are generated on the object display image; Obtaining the boundary sizes of the preset number of candidate regions, and normalizing the size of each candidate region to obtain a normalized candidate region; Classify each normalized candidate region to determine a target candidate region representing the target object; The boundary of the target candidate area is corrected and adjusted, and the image features of the corrected target candidate area are extracted as the image features of the target object.

3. The feature data generating method according to claim 2, in, The selective search algorithm is used to generate a preset number of candidate regions on the item display image, including: generating a plurality of initial image regions on the item display image; According to the region similarity and the first threshold, adjacent initial image regions are merged until all initial image regions are generated to obtain multiple initial segmented regions; According to the similarity between the initial segmented regions, the multiple initial segmented regions are merged until a preset number of candidate regions are obtained.

4. The feature data generating method according to claim 1, in, The performing attribute entity recognition according to the page description text data and obtaining the attribute characteristics of the target item based on the recognized attribute data includes: Input the page description text data into a pre-trained attribute extraction model to extract text attribute entities, wherein the sample annotation attribute entities in the sample data used to train the attribute extraction model are obtained by labeling according to the attribute parameters in the standard item library; According to the standard object library, the text attribute entity is standardized and cleaned to obtain the attribute characteristics of the target object; and The step of performing vector conversion according to the page description text data to obtain the text vector features of the target item includes: The word vector obtained by converting the text attribute entity vector is used as the text vector feature of the target item.

5. The feature data generating method according to claim 4, in, The step of performing standardized cleaning on the text attribute entity according to the standard object library includes: According to the basic attributes of the target item, a standard sub-item library matching the basic attributes is screened out from the standard item library, wherein the basic attributes include a brand logo and / or an item category; The text attribute entity is standardized and cleaned according to the standard sub-item library.

6. The feature data generating method according to claim 5, in, The step of performing standardized cleaning on the text attribute entity according to the standard sub-item library includes: For each attribute in the text attribute entity, determining the similarity between the extracted attribute value and the corresponding standard attribute values ​​in the standard sub-item library; According to the attribute similarity and the second threshold, it is determined whether to replace the extracted attribute value with the corresponding standard attribute value.

7. The feature data generating method according to claim 5, in, The step of performing standardized cleaning on the text attribute entity according to the standard object library further includes: Performing consistency check on the cleaned text attribute entity and the attribute parameters of the target item in the standard item library, so that both attribute data represent the same item; The verified text attribute entities are used as attribute features of the target item.

8. The feature data generating method according to claim 1, in, The step of fusing the image features, the text vector features and the attribute features to generate multi-feature data of the target object includes: The image features are converted into one-dimensional image features, and the one-dimensional image features, the text vector features and the attribute features are concatenated to obtain multi-feature data of the target object.

9. A feature data generating device, include: An acquisition unit is configured to acquire relevant data representing a target item on the platform, wherein the relevant data includes an item display image and page description text data; An image feature extraction unit is configured to extract features based on the object display image to obtain image features representing the target object; A text feature conversion unit is configured to perform vector conversion according to the page description text data to obtain a text vector feature of the target item; an attribute feature recognition unit, configured to perform attribute entity recognition according to the page description text data, and obtain attribute features of the target item based on the recognized attribute data; The feature fusion unit is configured to perform feature fusion on the image feature, the text vector feature and the attribute feature to generate multi-feature data of the target object.

10. A method for recommending items. include: In response to detecting an item search request, acquiring, from an item feature library, multi-feature data of a first item indicated by the item search request, wherein the multi-feature data in the item feature library is obtained by using the feature data generation method according to any one of claims 1 to 8; Determining the similarity between the multi-feature data of the first object and the multi-feature data of other objects in the object feature library, and determining the other objects with similarity greater than a third threshold as candidate objects for the first object; Generate item recommendation information based on the information of the first item and the candidate item for display.

11. The item recommendation method according to claim 10, wherein, generating the item recommendation information based on the information of the first item and the candidate item includes: performing consistency verification on the attribute features of the candidate item according to the attribute features of the first item to determine whether the attribute features of the candidate item are within the range of attribute feature changes of the first item; determining the candidate item with consistent verification as the same-item item of the first item, and using the information of the first item and the same-item item as the recommendation information.

12. An electronic device, comprising: one or more processors; a storage device having stored thereon one or more programs, when the one or more programs are executed by the one or more processors, causing the one or more processors to implement the method according to any one of claims 1-8 or 10-11.

13. A computer-readable medium having stored thereon a computer program, wherein, when the computer program is executed by a processor, the method according to any one of claims 1-8 or 10-11 is implemented.

14. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-8 or 10-11.