A labeling classification method for e-commerce products

Through the label classification method of e-commerce products, product matching and feature analysis are carried out based on keyword similarity and scene difference, and standard labels are generated, which solves the problem of low accuracy of product label recognition on e-commerce platforms and achieves higher accuracy of label recognition.

CN119474958BActive Publication Date: 2025-08-08BEIJING XINGTU DIGITAL NETWORK TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411352668.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-26
Publication Date
2025-08-08
Estimated Expiration
2044-09-26

AI Technical Summary

Technical Problem

In the prior art, the identification accuracy of e-commerce product labels is poor and cannot be classified according to the product usage scenarios.

Method used

By matching products based on the keyword similarity and fit of the target product, combining scene keyword differences, product image characteristics and background characteristics analysis, video frames are selectively supplemented, standard labels are generated or autonomous early warning processing is performed, and the label generation method is determined.

Benefits of technology

It improves the recognition accuracy of product labels, adapts to different application scenarios, avoids classification inaccurate problems caused by scene differences, and ensures that the label generation meets actual needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119474958B_ABST
    Figure CN119474958B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of label classification, and in particular to a label classification method for e-commerce products, comprising: matching target products according to keyword similarity and keyword compatibility to obtain a set of pre-matched products; determining a product processing method based on scene keyword differences to analyze product features or background features in product images of each pre-matched product in the pre-matched product set; determining image status based on the total number of product images and the proportion of valid images; selecting video frames for image supplementation based on product feature similarity or background feature similarity; determining whether to generate product labels based on keywords or perform autonomous warning processing by users based on background feature similarity or product feature similarity; and determining a standard label generation method based on the proportion of the maximum number of identical product labels. The present invention can classify products according to product usage scenarios and improve the recognition accuracy of product labels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of label classification, and in particular to a label classification method for e-commerce products. Background Art

[0002] In today's e-commerce environment, due to the non-standardization of product information, the same product may be used in different application scenarios. When users search by product name, it is difficult to accurately match relevant products that meet user needs, resulting in a poor user experience. Therefore, how to classify products according to product usage scenarios to improve the recognition accuracy of product labels is a technical problem that needs to be urgently solved by technical personnel in this field.

[0003] Chinese Patent Publication No. CN108491490A discloses a product labeling and identification system for e-commerce platforms. The system includes: a user behavior information acquisition module for acquiring user browsing and consumption information on the e-commerce platform; a user behavior information classification and integration module for classifying and integrating the acquired user behavior information; a user behavior information analysis module for analyzing the integrated user behavior information; and a comparison module for comparing user behavior information with product labels. This technical solution has the following problems: It fails to classify products based on their usage scenarios, resulting in poor recognition accuracy of product labels. Summary of the Invention

[0004] To this end, the present invention provides a labeling classification method for e-commerce products to overcome the problem in the prior art that products are not classified according to product usage scenarios, resulting in poor recognition accuracy of product labels.

[0005] To achieve the above-mentioned object, the present invention provides a labeling classification method for e-commerce products, comprising: matching target products based on keyword similarity and keyword compatibility of product texts corresponding to the target products to obtain a set of pre-matched products;

[0006] Detect the difference between the scene keywords corresponding to the pre-matched product set, and determine the product processing method based on the difference between the scene keywords;

[0007] The product processing method is to analyze the product features or background features in the product image of each pre-matched product in the pre-matched product set;

[0008] Determine the image status based on the total number of product images and the percentage of valid images;

[0009] When the total number of product images is less than the preset total number of product images or the ratio of valid images is less than the preset ratio of valid images, video frames are selected for image supplementation based on the similarity of product features or background features;

[0010] Generate product tags based on keywords or perform autonomous warning processing for users based on the similarity of background features or product features;

[0011] The method for generating standard tags is determined based on the maximum ratio of the number of identical product tags, which is to convert product tags into standard tags or to use keyword combinations as standard tags.

[0012] Furthermore, product matching includes:

[0013] When matching a single target product, the product matching status is determined based on the keyword similarity and keyword compatibility between the target product and the reference product. The reference products and the target product that are in the matching status are collectively recorded as a pre-matched product set.

[0014] When the pre-matched product set is established, the target products and reference products in the set are recorded as products to be standardized.

[0015] Furthermore, the effective image ratio is determined based on the ratio of the number of effective images to the total number of product images.

[0016] Furthermore, when analyzing the product features in the product images of each to-be-standardized product in the pre-matched product set, if the image status is that the total number of product images is less than the preset total number of product images or the proportion of valid images is less than the preset proportion of valid images, video frames are selected for image supplementation based on the similarity of product features:

[0017] Record the video frames corresponding to the reference videos whose product feature similarity is greater than the preset product feature similarity as the first-class video frames, record the first-class video frames with the largest product feature similarity as the target first-class video frames, record the other first-class video frames excluding the target first-class video frames as the reference first-class video frames, perform contour difference detection on the target first-class video frames, record the target first-class video frames and the reference video frames with the largest contour difference from the target first-class video frames as a first-class combination, continue to perform difference detection on the first-class combination, add the reference first-class video frame with the largest average contour difference from each first-class video frame in the first-class combination to the first-class combination, until the number of first-class video frames in the first-class combination is equal to the pre-extracted number;

[0018] The reference videos include product review videos and product videos.

[0019] Furthermore, the contour difference is determined according to the absolute value of the difference between the contour length corresponding to the target type of video frame and the contour length corresponding to the reference type of video frame.

[0020] Furthermore, when analyzing the background features of the product images of each to-be-standardized product in the pre-matched product set, if the image status is that the total number of product images is less than the preset total number of product images or the proportion of valid images is less than the preset proportion of valid images, video frames are selected for image supplementation based on the similarity of background features:

[0021] Recording the video frames corresponding to the reference video whose background feature similarity is greater than the preset background feature similarity as the second-category video frames, sorting the second-category video frames in descending order of background area to obtain a second-category video frame sequence, and extracting a number of second-category video frames from each frame extraction point in the second-category video frame sequence;

[0022] The frame extraction points are evenly distributed in the second type of video frame sequence and the number of the frame extraction points is determined according to the pre-extraction number.

[0023] Furthermore, the product feature similarity is determined based on the ViT product feature value corresponding to the product image and the ViT product feature value corresponding to the video frame;

[0024] The background feature similarity is determined based on the ViT background feature value corresponding to the product image and the ViT background feature value corresponding to the video frame.

[0025] Furthermore, if the scene keyword difference is less than the preset scene keyword difference, the background features in the product images of each to-be-standardized product in the pre-matched product set are analyzed:

[0026] Perform background feature similarity detection on the products to be standardized, and determine the generation of product tags based on keywords or user autonomous warning processing based on the background feature similarity;

[0027] If the scene keyword difference is greater than or equal to the preset scene keyword difference, analyze the product features in the product images of each product to be standardized in the pre-matched product set:

[0028] Carry out product feature similarity detection for products to be standardized, and determine whether to generate product tags based on keywords or perform autonomous warning processing for users based on the similarity of product features.

[0029] Furthermore, if the maximum number of identical product labels accounts for a greater than or equal to a preset maximum number of identical product labels, the product labels corresponding to the maximum number of identical product labels are converted into standard labels;

[0030] If the maximum ratio of the number of identical product tags is less than the preset maximum ratio of the number of identical product tags, identify the identical keywords in each product tag and convert the keyword combination consisting of the identical keywords into a standard tag.

[0031] Furthermore, the proportion of the maximum number of identical product labels is determined based on the ratio of the maximum number of identical product labels to the total number of products to be standardized.

[0032] Compared with the prior art, the beneficial effect of the present invention lies in that, in the technical solution of the present invention, product matching is performed for target products based on keyword similarity and keyword compatibility, and the keyword similarity and keyword compatibility effectively reflect the similarity and compatibility of the target product keywords, and then the target products are classified to obtain a pre-matched product set, so that the correlation degree of each target product in the pre-matched product set is relatively high, which is conducive to the subsequent determination of the product processing method based on the difference in scene keywords.

[0033] Furthermore, the present invention determines the commodity processing method based on the difference in scene keywords. The difference in scene keywords effectively reflects the difference in usage scenarios of each target commodity in the pre-matched commodity set, and then adaptively selects different commodity processing methods, so that the selection of commodity processing method is more in line with the actual application scenario, avoiding the problem that the target commodity cannot be accurately classified due to different usage scenarios, thereby improving the recognition accuracy of commodity labels.

[0034] Furthermore, the present invention determines the image status based on the total number of product images and the proportion of valid images. The total number of product images and the proportion of valid images reflect the quality of the product images, and then determines whether it is necessary to select video frames for image supplementation based on the similarity of product features or background features according to the actual application scenario. This avoids the problem of poor accuracy in analysis based on product features or background features in product images due to lack of featureness of product images, thereby accurately classifying products and improving the recognition accuracy of product labels.

[0035] Furthermore, the present invention effectively reflects the similarity of the products to be standardized through the similarity of background features and product features, and then selects to generate product labels based on keywords or users to perform autonomous early warning processing according to the actual application scenario, thereby avoiding the problem of poor product similarity after establishing standard labels for each target product, thereby improving the recognition accuracy of product labels.

[0036] Furthermore, the present invention effectively reflects the number of the maximum number of identical product labels through the proportion of the maximum number of identical product labels, and then adaptively selects the standard label generation method, so that the selection of the label generation method is more in line with the actual application scenario, avoiding the problem of poor representativeness of the standard label due to the small number of identical labels, thereby improving the recognition accuracy of the product labels. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 A schematic diagram of a labeling classification method for e-commerce products according to the present invention;

[0038] Figure 2 This is a flow chart of the present invention for determining a product processing method based on the difference between scene keywords;

[0039] Figure 3 This is a flow chart of the present invention for determining, based on image status, whether to select a video frame for image supplementation or not based on product feature similarity or background feature similarity;

[0040] Figure 4 This is a flowchart of the present invention determining that a standard tag generation method is to convert a product tag into a standard tag or to use a keyword combination as a standard tag according to the maximum proportion of the number of identical product tags. DETAILED DESCRIPTION

[0041] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.

[0042] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0043] It should be noted that, in the description of the present invention, terms such as "up", "down", "left", "right", "inside", and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present invention.

[0044] Furthermore, it should be noted that, in the description of the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.

[0045] See also Figures 1 to 4 As shown, the present invention provides a labeling classification method for e-commerce products, including:

[0046] Match the target product based on the keyword similarity and keyword fit of the product text corresponding to the target product to obtain a pre-matched product set;

[0047] Detect the difference between the scene keywords corresponding to the pre-matched product set, and determine the product processing method based on the difference between the scene keywords;

[0048] The product processing method is to analyze the product features or background features in the product image of each pre-matched product in the pre-matched product set;

[0049] Determine the image status based on the total number of product images and the percentage of valid images;

[0050] When the total number of product images is less than the preset total number of product images or the ratio of valid images is less than the preset ratio of valid images, video frames are selected for image supplementation based on the similarity of product features or background features;

[0051] Generate product tags based on keywords or perform autonomous warning processing for users based on the similarity of background features or product features;

[0052] The method for generating standard tags is determined based on the maximum ratio of the number of identical product tags, which is to convert product tags into standard tags or to use keyword combinations as standard tags.

[0053] The application scenario of the present invention is the labeling processing of product information on an e-commerce platform. The target products include product pictures, product videos, product texts and product review videos. Product pictures are pictures uploaded by merchants to the e-commerce platform containing the appearance and usage scenarios of the target products. Product videos are videos uploaded by merchants to the e-commerce platform containing the product usage process, product function demonstration or appearance display, etc. Product text is a character paragraph containing information such as the product name, product specifications and product use. Product review videos are videos uploaded by consumers to the e-commerce platform after purchasing the products. This is content that is easy for technicians in this field to understand and will not be elaborated here.

[0054] If the scene keyword difference is greater than or equal to the preset scene keyword difference, the product processing method is to analyze the product features in the product image of each product to be standardized in the pre-matched product set;

[0055] If the scene keyword difference is less than the preset scene keyword difference, the product processing method is to analyze the background features in the product image of each product to be standardized in the pre-matched product set;

[0056] Image status includes: the total number of product images is less than the preset total number of product images or the ratio of valid images is less than the preset ratio of valid images; and the total number of product images is greater than or equal to the preset total number of product images and the ratio of valid images is greater than or equal to the preset ratio of valid images;

[0057] When the total number of product images is less than the preset total number of product images or the ratio of valid images is less than the preset ratio of valid images, video frames are selected for image supplementation based on product feature similarity or background feature similarity, and product features or background features in the supplemented video frames and the product images of each pre-matched product in the pre-matched product set are analyzed;

[0058] When the total number of product images is greater than or equal to the preset total number of product images and the percentage of valid images is greater than or equal to the preset percentage of valid images, image supplementation is not required, and the product features or background features in the product images of each pre-matched product in the pre-matched product set are analyzed;

[0059] Scene keywords include but are not limited to living room, bedroom, hiking, camping, rock climbing, mountain biking, Christmas, Valentine's Day, workplace and conference room. The recognition of scene keywords can be achieved through machine vision and deep learning networks. This is content that is easy for technical personnel in this field to understand and will not be elaborated here.

[0060] Scenario keyword difference = (total number of products in the collection - number of target products corresponding to the maximum same scenario keyword) / total number of products in the collection. The total number of products in the collection is the number of products to be standardized in a single pre-matched product collection. The method for confirming the maximum same scenario keyword is to perform the same scenario keyword detection on each product to be standardized in the pre-matched product collection. When performing the same scenario keyword detection on a single product to be standardized, the product to be standardized is recorded as the target product to be standardized, and other products to be standardized in the pre-matched product collection that do not include the target product to be standardized are recorded as reference products to be standardized. The reference products to be standardized and the target products to be standardized that have the same scenario keywords as the target products to be standardized are recorded as the same scenario keyword combination, and the same scenario keyword detection is performed on the products to be standardized that are not recorded as the same scenario keyword combination, until all products to be standardized are recorded as the same scenario keyword combination, and the number of products to be standardized corresponding to the same scenario keyword combination that contains the largest number of products to be standardized is recorded as the target number of products corresponding to the maximum same scenario keyword.

[0061] The value of the preset scene keyword difference can be determined by the user according to the actual application scenario. The higher the user's demand for the accuracy of setting the standard label, the smaller the value of the preset scene keyword difference. A value of the preset scene keyword difference is provided, and the preset scene keyword difference is 20%.

[0062] The total number of product images is the number of product images corresponding to a single target product. The values of the preset total number of product images and the preset ratio of valid images can be determined by the user according to the actual application scenario. The higher the user's demand for the accuracy of setting standard labels, the larger the values of the preset total number of product images and the preset ratio of valid images. The values of the preset total number of product images and the preset ratio of valid images are provided. The preset total number of product images is the average value of the product images corresponding to each target product in the e-commerce platform, and the preset ratio of valid images is 80%.

[0063] Specifically, product matching includes:

[0064] When matching a single target product, the product matching status is determined based on the keyword similarity and keyword compatibility between the target product and the reference product. The reference products and the target product that are in the matching status are collectively recorded as a pre-matched product set.

[0065] When the pre-matched product set is established, the target products and reference products in the set are recorded as products to be standardized.

[0066] Keyword similarity = number of identical keywords / number of reference keywords. The number of reference keywords is determined as follows: if the number of target product keywords is greater than or equal to the number of reference product keywords, the number of reference keywords is the target product keyword number; if the number of target product keywords is less than the reference product keyword number, the number of reference keywords is the reference product keyword number. The target product keyword number is the number of keywords identified in the product text corresponding to the target product, while the reference product keyword number is the number of keywords identified in the product text corresponding to the reference product.

[0067] Keyword fit = number of matching keywords / (number of reference product keywords + number of target product keywords). The number of matching keywords is the number of keywords in which the proportion of the number of times a single keyword corresponding to the target product is used in combination with the reference product is greater than the proportion of the preset number of times the combination is used. The proportion of the number of times the combination is used = the number of times two keywords appear simultaneously in the product text corresponding to the target product on the e-commerce platform / the number of target products on the e-commerce platform.

[0068] The matching status of the matching product is that the keyword similarity is greater than the preset keyword similarity and the keyword fit is greater than the preset keyword fit. The values of the preset keyword similarity and the preset keyword fit can be determined by the user according to the actual application scenario. The higher the user's demand for the accuracy of setting the standard label, the greater the values of the preset keyword similarity and the preset keyword fit. A value of the preset keyword similarity and the preset keyword fit is provided, and the preset keyword similarity is 70%, and the preset keyword fit is 70%.

[0069] Specifically, the effective image ratio is determined based on the ratio of the number of effective images to the total number of product images.

[0070] Valid image ratio = number of valid images / total number of product images. Valid images are product images in which both the product area ratio and the background area ratio are greater than the preset area ratios. Product area ratio = product area / product image area. Background area ratio = background area / product image area. The product area is the area of the target product in the product image. The background area and product area are segmented using edge detection technology. The product area is calculated using contourArea in OpenCV. This is easily understood by those skilled in the art and is not elaborated here. Background area = product image area - product area. Product image area = product image length × product image width.

[0071] Specifically, when analyzing the product features in the product images of each to-be-standardized product in the pre-matched product set, if the image status is that the total number of product images is less than the preset total number of product images or the proportion of valid images is less than the preset proportion of valid images, video frames are selected for image supplementation based on the similarity of product features:

[0072] Record the video frames corresponding to the reference videos whose product feature similarity is greater than the preset product feature similarity as the first-class video frames, record the first-class video frames with the largest product feature similarity as the target first-class video frames, record the other first-class video frames excluding the target first-class video frames as the reference first-class video frames, perform contour difference detection on the target first-class video frames, record the target first-class video frames and the reference video frames with the largest contour difference from the target first-class video frames as a first-class combination, continue to perform difference detection on the first-class combination, add the reference first-class video frame with the largest average contour difference from each first-class video frame in the first-class combination to the first-class combination, until the number of first-class video frames in the first-class combination is equal to the pre-extracted number;

[0073] The reference videos include product review videos and product videos.

[0074] The value of the preset product feature similarity can be determined by the user according to the actual application scenario. The higher the user's demand for the accuracy of setting the standard label, the greater the value of the preset product feature similarity. A value of the preset product feature similarity is provided, and the preset product feature similarity is 0.8.

[0075] Pre-extracted quantity = total number of preset product images - number of product images. The calculation formula for the average value μ of the contour difference is: Where N is the number of video frames of a class in a class combination, i is the i-th video frame of a class in a class combination, i = 1, 2, 3, ..., N, x i is the contour difference of the i-th video frame in the one-class combination.

[0076] Specifically, the contour difference is determined according to the absolute value of the difference between the contour length corresponding to the target type video frame and the contour length corresponding to the reference type video frame.

[0077] The contour length is the length of the dividing line between the product area and the background area. The contour length is calculated using arcLength in OpenCV. This is well understood by those skilled in the art and will not be elaborated on here.

[0078] Specifically, when analyzing the background features of the product images of each to-be-standardized product in the pre-matched product set, if the image status is that the total number of product images is less than the preset total number of product images or the proportion of valid images is less than the preset proportion of valid images, video frames are selected for image supplementation based on the similarity of background features:

[0079] Recording the video frames corresponding to the reference video whose background feature similarity is greater than the preset background feature similarity as the second-category video frames, sorting the second-category video frames in descending order of background area to obtain a second-category video frame sequence, and extracting a number of second-category video frames from each frame extraction point in the second-category video frame sequence;

[0080] The frame extraction points are evenly distributed in the second type of video frame sequence and the number of the frame extraction points is determined according to the pre-extraction number.

[0081] The frame extraction point is the position where the video frame is extracted from the second-category video frame sequence. If the number of frame extraction points is m, the second-category video frame sequence is divided into (m-1) equal parts, and the frame extraction points are located at the starting end, the end end, and the positions of the equal division points of the second-category video frame sequence. The number of frame extraction points m is equal to the pre-extraction number.

[0082] The value of the preset background feature similarity can be determined by the user according to the actual application scenario. The higher the user's demand for the accuracy of setting the standard label, the larger the value of the preset background feature similarity. A value of the preset background feature similarity is provided, and the preset background feature similarity is 0.8.

[0083] Specifically, the product feature similarity is determined based on the ViT product feature value corresponding to the product image and the ViT product feature value corresponding to the video frame;

[0084] The background feature similarity is determined based on the ViT background feature value corresponding to the product image and the ViT background feature value corresponding to the video frame.

[0085] The method for confirming the ViT product feature value is to divide the product area into several image blocks of the same area through the ViT model, and input the image blocks into the linear projection layer to obtain a vector group recorded as the ViT product feature value; the method for confirming the ViT background feature value is to divide the background area into several image blocks of the same area through the ViT model, and input the image blocks into the linear projection layer to obtain a vector group recorded as the ViT background feature value.

[0086] Both product feature similarity and background feature similarity are calculated using cosine similarity. The cosine value of the angle between the ViT product feature value corresponding to the product image and the ViT product feature value corresponding to the video frame is recorded as the product feature similarity, and the cosine value of the angle between the ViT background feature value corresponding to the product image and the ViT background feature value corresponding to the video frame is recorded as the background feature similarity. The value range of product feature similarity and background feature similarity is [0,1].

[0087] Specifically, if the scene keyword difference is less than the preset scene keyword difference, the background features in the product images of each to-be-standardized product in the pre-matched product set are analyzed:

[0088] Perform background feature similarity detection on the products to be standardized, and determine the generation of product tags based on keywords or user autonomous warning processing based on the background feature similarity;

[0089] If the scene keyword difference is greater than or equal to the preset scene keyword difference, analyze the product features in the product images of each product to be standardized in the pre-matched product set:

[0090] Carry out product feature similarity detection for products to be standardized, and determine whether to generate product tags based on keywords or perform autonomous warning processing for users based on the similarity of product features.

[0091] Perform background feature similarity detection on the products to be standardized. If the background feature similarity is greater than or equal to the preset background feature similarity, generate product labels based on keywords; if the background feature similarity is less than the preset background feature similarity, the user will perform autonomous warning processing;

[0092] The product feature similarity test is performed on the products to be standardized. If the product feature similarity is greater than or equal to the preset product feature similarity, a product label is generated based on the keyword; if the product feature similarity is less than the preset product feature similarity, the user will perform an independent warning process;

[0093] Generating product labels based on keywords includes: converting each keyword in the product text corresponding to the product to be standardized into a product label;

[0094] User-initiated alert processing includes manual review to convert product labels into standard labels.

[0095] Specifically, if the maximum number of identical product labels accounts for a greater than or equal to the preset maximum number of identical product labels, the product labels corresponding to the maximum number of identical product labels are converted into standard labels;

[0096] If the maximum ratio of the number of identical product tags is less than the preset maximum ratio of the number of identical product tags, identify the identical keywords in each product tag and convert the keyword combination consisting of the identical keywords into a standard tag.

[0097] The value of the preset maximum ratio of the number of identical product labels can be determined by the user according to the actual application scenario. The higher the user's demand for the accuracy of setting standard labels, the larger the value of the preset maximum ratio of the number of identical product labels. A value of the preset maximum ratio of the number of identical product labels is provided, and the preset maximum ratio of the number of identical product labels is 80%;

[0098] Converting the product tags corresponding to the maximum number of identical product tags into standard tags includes: it can be understood that the order of keywords in the product tags corresponding to the maximum number of identical product tags may be different, and randomly selecting a product tag from the product tags corresponding to the maximum number of identical product tags as the standard tag for the pre-matched product set.

[0099] Specifically, the proportion of the maximum number of identical product labels is determined based on the ratio of the maximum number of identical product labels to the total number of products to be standardized.

[0100] The maximum proportion of identical product labels = the maximum number of identical product labels / the total number of products to be standardized in the pre-matched product set. The total number of products to be standardized in the pre-matched product set is the number of products to be standardized contained in the pre-matched product set. The maximum number of identical product labels is confirmed by performing identical label detection on the product labels corresponding to each product to be standardized in the pre-matched product set. When performing identical label detection on the product label corresponding to a single product to be standardized, the product label corresponding to the product to be standardized is recorded as the target label, and the product labels corresponding to the products to be standardized that do not contain the target label in the pre-matched product set are recorded as reference labels. The reference label and the target label that are identical to the keywords corresponding to the target label are recorded as the identical product label combination, and identical label detection is performed on the reference labels that are not recorded as the identical product label combination until all product labels are recorded as the identical product label combination, and the number of product labels corresponding to the identical product label combination containing the largest number of product labels is recorded as the maximum number of identical product labels.

[0101] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.

[0102] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A labeling classification method for e-commerce products, characterized by: include: Match the target product based on the keyword similarity and keyword fit of the product text corresponding to the target product to obtain a pre-matched product set; Detect the difference between the scene keywords corresponding to the pre-matched product set, and determine the product processing method based on the difference between the scene keywords; The product processing method is to analyze the product features or background features in the product image of each pre-matched product in the pre-matched product set; Determine the image status based on the total number of product images and the percentage of valid images; When the total number of product images is less than the preset total number of product images or the ratio of valid images is less than the preset ratio of valid images, video frames are selected for image supplementation based on the similarity of product features or background features; Generate product tags based on keywords or perform autonomous warning processing for users based on the similarity of background features or product features; The method for generating standard tags is determined based on the maximum ratio of the number of identical product tags, which is to convert product tags into standard tags or to use keyword combinations as standard tags.

2. The labeling and classification method for e-commerce products according to claim 1 is characterized in that: Product matching includes: When matching a single target product, the product matching status is determined based on the keyword similarity and keyword compatibility between the target product and the reference product. The reference products and the target product that are in the matching status are collectively recorded as a pre-matched product set. When the pre-matched product set is established, the target products and reference products in the set are recorded as products to be standardized.

3. The labeling classification method for e-commerce products according to claim 2 is characterized in that: The effective image ratio is determined based on the ratio of the number of effective images to the total number of product images.

4. The labeling and classification method for e-commerce products according to claim 3, characterized in that: When analyzing the product features in the product images of each to-be-standardized product in the pre-matched product set, if the image status is that the total number of product images is less than the preset total number of product images or the proportion of valid images is less than the preset proportion of valid images, video frames are selected for image supplementation based on the similarity of product features: Record the video frames corresponding to the reference videos whose product feature similarity is greater than the preset product feature similarity as the first-class video frames, record the first-class video frames with the largest product feature similarity as the target first-class video frames, record the other first-class video frames excluding the target first-class video frames as the reference first-class video frames, perform contour difference detection on the target first-class video frames, record the target first-class video frames and the reference video frames with the largest contour difference from the target first-class video frames as a first-class combination, continue to perform difference detection on the first-class combination, add the reference first-class video frame with the largest average contour difference from each first-class video frame in the first-class combination to the first-class combination, until the number of first-class video frames in the first-class combination is equal to the pre-extracted number; The reference videos include product review videos and product videos.

5. The labeling classification method for e-commerce products according to claim 4 is characterized in that: The contour difference is determined according to the absolute value of the difference between the contour length corresponding to the target type video frame and the contour length corresponding to the reference type video frame.

6. The labeling classification method for e-commerce products according to claim 5 is characterized in that: When analyzing the background features of the product images of each to-be-standardized product in the pre-matched product set, if the total number of product images is less than the preset total number of product images or the proportion of valid images is less than the preset proportion of valid images, video frames are selected for image supplementation based on the similarity of background features: Recording the video frames corresponding to the reference video whose background feature similarity is greater than the preset background feature similarity as the second-category video frames, sorting the second-category video frames in descending order of background area to obtain a second-category video frame sequence, and extracting a number of second-category video frames from each frame extraction point in the second-category video frame sequence; The frame extraction points are evenly distributed in the second type of video frame sequence and the number of the frame extraction points is determined according to the pre-extraction number.

7. The labeling and classification method for e-commerce products according to claim 6, characterized in that: The product feature similarity is determined based on the ViT product feature value corresponding to the product image and the ViT product feature value corresponding to the video frame; The background feature similarity is determined based on the ViT background feature value corresponding to the product image and the ViT background feature value corresponding to the video frame.

8. The labeling and classification method for e-commerce products according to claim 7, characterized in that: If the scene keyword difference is less than the preset scene keyword difference, analyze the background features in the product images of each to-be-standardized product in the pre-matched product set: Perform background feature similarity detection on the products to be standardized, and determine the generation of product tags based on keywords or user autonomous warning processing based on the background feature similarity; If the scene keyword difference is greater than or equal to the preset scene keyword difference, analyze the product features in the product images of each product to be standardized in the pre-matched product set: Carry out product feature similarity detection for products to be standardized, and determine whether to generate product tags based on keywords or perform autonomous warning processing for users based on the similarity of product features.

9. The labeling classification method for e-commerce products according to claim 8, characterized in that: If the maximum number of identical product labels is greater than or equal to the preset maximum number of identical product labels, the product labels corresponding to the maximum number of identical product labels are converted into standard labels; If the maximum ratio of the number of identical product tags is less than the preset maximum ratio of the number of identical product tags, identify the identical keywords in each product tag and convert the keyword combination consisting of the identical keywords into a standard tag.

10. The labeling classification method for e-commerce products according to claim 9, characterized in that: The proportion of the maximum number of identical product labels is determined based on the ratio of the maximum number of identical product labels to the total number of products to be standardized.

Citation Information

Patent Citations

  • Labeling differentiation and recognition system for e-commerce platform commodities and method thereof

    CN108491490A

  • E-commerce platform commodity matching method and device and readable storage medium

    CN110083678A

  • Automatic labeling method for quality problem scene labels based on categories

    CN112579776A