Category classification method and device, electronic equipment and storage medium

By acquiring video data to generate product feature vectors and using a target category prediction table and a deep learning model to determine product hierarchical categories, this solves the problem of low efficiency and low accuracy of manual review on e-commerce platforms, and achieves efficient and accurate automatic identification of product categories.

CN114429599BActive Publication Date: 2026-05-12BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
Filing Date
2021-12-24
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

The manual review of product categories on existing e-commerce platforms is inefficient and inaccurate.

Method used

By acquiring video data, product feature vectors for product information are generated. Multiple target category prediction tables are used to determine the target hierarchical category of product information based on a tree-like relationship. Deep learning models and Beam Search methods are employed to improve the accuracy of category classification.

Benefits of technology

It improved the accuracy of category classification and achieved more efficient automatic identification of product categories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114429599B_ABST
    Figure CN114429599B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a category classification method and device, electronic equipment and storage medium, and relates to the technical field of computers. The method comprises: obtaining video data by implementing the category classification method provided in the embodiments of the present disclosure; wherein the video data comprises commodity information; processing the video data to generate a commodity feature vector of the commodity information; determining a target hierarchical category corresponding to the commodity information according to the commodity feature vector and a plurality of target category prediction tables; wherein each target category prediction table comprises a plurality of category vectors, each target category prediction table corresponds to a category level, and the category vectors in different target category prediction tables have a tree relationship. Thus, the accuracy of category classification can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a category classification method, apparatus, electronic device, and storage medium. Background Technology

[0002] An e-commerce product database is the core foundation of e-commerce sales. All major existing internet e-commerce platforms need to establish their own product databases, publishing information on available products for sale daily, allowing users to select and transact. During the product review and publication phase, the platform needs to review the products published by merchants, including verifying whether the product categories are correctly matched.

[0003] In related technologies, platforms identify product categories in merchant-published product information through manual review, but this method is inefficient and inaccurate. Summary of the Invention

[0004] This disclosure provides a category classification method, apparatus, electronic device, and storage medium to at least address the problems of low efficiency and low accuracy of manual review in e-commerce platforms in related technologies. The technical solution of this disclosure is as follows:

[0005] According to a first aspect of the present disclosure, a category classification method is provided, comprising: acquiring video data; wherein the video data includes product information; processing the video data to generate product feature vectors of the product information; and determining the target hierarchical category corresponding to the product information based on the product feature vectors and multiple target category prediction tables; wherein each target category prediction table includes multiple category vectors, each target category prediction table corresponds to a category level, and the category vectors in different target category prediction tables have a tree-like relationship. This improves the accuracy of category classification.

[0006] According to a second aspect of the present disclosure, a category classification apparatus is provided, comprising: a data acquisition unit for acquiring video data, wherein the video data includes product information; a data processing unit for processing the video data to generate a product feature vector of the product information; and a category determination unit for determining a target level category corresponding to the product information based on the product feature vector and multiple target category prediction tables; wherein each target category prediction table includes multiple category vectors, each target category prediction table corresponds to a category level, and the category vectors in different target category prediction tables have a tree-like relationship.

[0007] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the category classification method as described in the first aspect above.

[0008] According to a fourth aspect of the present disclosure, a storage medium is provided that, when instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform the category classification method as described in the first aspect above.

[0009] According to a fifth aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the category classification method as described in the first aspect above.

[0010] The technical solutions provided by the embodiments of this disclosure bring at least the following beneficial effects:

[0011] The category classification method provided in this embodiment acquires video data, including product information. The video data is processed to generate product feature vectors for the product information. Based on the product feature vectors and multiple target category prediction tables, the target category level corresponding to the product information is determined. Each target category prediction table includes multiple category vectors, and each target category prediction table corresponds to a category level. The category vectors in different target category prediction tables have a tree-like relationship. This improves the accuracy of category classification.

[0012] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0013] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0014] Figure 1 This is a flowchart illustrating a category classification method according to an exemplary embodiment;

[0015] Figure 2 This is a flowchart illustrating another category classification method according to an exemplary embodiment;

[0016] Figure 3 This is a flowchart illustrating yet another category classification method according to an exemplary embodiment;

[0017] Figure 4 This is a flowchart illustrating yet another category classification method according to an exemplary embodiment;

[0018] Figure 5 This is a flowchart illustrating yet another category classification method according to an exemplary embodiment;

[0019] Figure 6 This is a structural diagram of a category classification device according to an exemplary embodiment;

[0020] Figure 7 This is a structural diagram of a category determination unit of a category classification device according to an exemplary embodiment;

[0021] Figure 8 This is a structural diagram of a data processing unit of a category classification device according to an exemplary embodiment;

[0022] Figure 9 This is a structural diagram of the data processing unit of another category classification apparatus according to an exemplary embodiment;

[0023] Figure 10 This is a structural diagram of another category classification device according to an exemplary embodiment;

[0024] Figure 11 This is a structural diagram of yet another category classification device according to an exemplary embodiment;

[0025] Figure 12 This is a structural diagram of an updating unit of a category classification device according to an exemplary embodiment;

[0026] Figure 13 This is a block diagram illustrating an electronic device according to an exemplary embodiment. Detailed Implementation

[0027] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0028] Unless the context otherwise requires, throughout the specification and claims, the term "comprising" is interpreted as open-ended and encompassing, meaning "including, but not limited to." In the description of the specification, terms such as "some embodiments" are intended to indicate that a particular feature, structure, material, or characteristic associated with that embodiment or example is included in at least one embodiment or example of this disclosure. The illustrative representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics mentioned may be included in any suitable manner in any one or more embodiments or examples.

[0029] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0030] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0031] It should be noted that the acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0032] It should be noted that the category classification method of this disclosure embodiment can be executed by the category classification device of this disclosure embodiment. The category classification device can be implemented by software and / or hardware, and can be configured in an electronic device, wherein the electronic device can install and run a category classification program. The electronic device can include, but is not limited to, hardware devices with various operating systems such as smartphones and tablets.

[0033] Figure 1 This is a flowchart illustrating a category classification method according to an exemplary embodiment.

[0034] like Figure 1 As shown, including but not limited to the following steps:

[0035] S1: Obtain video data; the video data includes product information.

[0036] With the continuous development of internet technology, more and more e-commerce platforms are offering live streaming for product sales. These platforms utilize video conferencing to broadcast product demonstrations and other content online, leveraging the internet's intuitiveness, speed, superior presentation, rich content, strong interactivity, and lack of geographical limitations to enhance product promotion. Furthermore, after the live stream concludes, replays and on-demand viewing are available at any time, effectively extending the duration and reach of the live stream and maximizing its value. However, with the increasing number of merchants joining e-commerce platforms and the diverse range of products displayed in different live streams and replays, it is necessary to categorize the products shown in the videos according to their categories to facilitate quick access for users.

[0037] In this embodiment of the disclosure, acquiring video data can be acquiring video data that is being broadcast live or replayed, or acquiring video data within a specific time period that is being broadcast live or replayed.

[0038] It is understood that video data includes target images and / or target text. Target images may include product information, such as images and text descriptions, while target text may include product information, such as audio information introducing the product.

[0039] S2: Process the video data to generate product feature vectors for product information.

[0040] In this embodiment of the disclosure, the acquired video data is processed to generate a product feature vector of product information.

[0041] Among these methods, video data can be processed using any deep learning model and encoded into product feature vectors of a specific length.

[0042] S3: Based on the product feature vector and multiple target category prediction tables, determine the target level category corresponding to the product information; where each target category prediction table includes multiple category vectors, each target category prediction table corresponds to a category level, and the category vectors in different target category prediction tables have a tree-like relationship.

[0043] It is understood that multiple category levels corresponding to multiple target category prediction tables may include two or more category levels, and this disclosure does not impose specific limitations on this. For example, multiple category levels may include three levels, such as: first-level categories including "clothing", second-level categories including "men's clothing", and third-level categories including "jeans".

[0044] There is a hierarchical tree relationship between different categories. For example, the first-level category is "clothing". The second-level categories corresponding to the first-level category "clothing" can be "men's clothing", "women's clothing", etc. The first-level category can correspond to one or more second-level categories, and the second-level categories can also correspond to one or more third-level categories, and so on, forming a hierarchical tree relationship.

[0045] In this embodiment of the disclosure, the categories in the category hierarchy are processed in advance to generate category vectors. Then, the generated category vectors are summarized to generate a target category prediction table for the corresponding level. The method of processing categories to generate category vectors in this embodiment of the disclosure can be the same as the method of processing video data to generate product feature vectors for product information.

[0046] For example, at least one category in the first-level category is generated into a category vector, and then the vectors are summarized to generate a target category prediction table corresponding to the first-level category; at least one category in the second-level category is generated into a category vector, and then the vectors are summarized to generate a target category prediction table corresponding to the second-level category, etc.

[0047] It is understood that in this embodiment of the disclosure, the target hierarchical category corresponding to the product information is determined based on the product feature vector and multiple target category prediction tables. The target hierarchical category may include category information of all levels. For example, when the category hierarchy includes four levels, the target hierarchical category corresponding to the product information is determined, including the corresponding first-level category, second-level category, third-level category and fourth-level category.

[0048] By implementing the category classification method provided in this embodiment, video data is acquired; the video data includes product information; the video data is processed to generate product feature vectors for the product information; based on the product feature vectors and multiple target category prediction tables, the target hierarchical category corresponding to the product information is determined; each target category prediction table includes multiple category vectors, each target category prediction table corresponds to a category level, and the category vectors in different target category prediction tables have a tree-like relationship. Therefore, by utilizing the tree-like relationship of hierarchical categories and drawing on the ranking and recall method to achieve category classification, the accuracy of category classification can be improved.

[0049] Figure 2 This is a flowchart illustrating another category classification method according to an exemplary embodiment.

[0050] like Figure 2 As shown, including but not limited to the following steps:

[0051] S21: Obtain video data; wherein, the video data includes product information.

[0052] The description of S21 in this embodiment can be found in the description of S1 in the above embodiments, and will not be repeated here.

[0053] S22: Process the video data to generate product feature vectors for product information.

[0054] In some embodiments, processing video data to generate a product feature vector of product information includes: extracting frames from the video data to obtain a target image; wherein the target image includes product information; and inputting the target image into a trained image processing model to generate a product feature vector of product information.

[0055] It is understandable that in the scenario of influencer live streaming e-commerce, the live video data or recorded live video data will contain images of the product's appearance and a brief description when introducing the product. It is also conceivable that only one product will be introduced at a time. By extracting frames from the video data, only one product's information will be contained in a single frame.

[0056] In this embodiment of the disclosure, after acquiring the target image, the target image is input into a trained image processing model. The image processing model can be a residual network and a Vision Transformer network model, thereby obtaining the product feature vector of the product information included in the target image.

[0057] In other embodiments, the video data is processed to generate a product feature vector of product information, including: performing speech recognition on the video data to obtain target text; wherein the target text includes product information; and inputting the target text into a trained deep language representation model to generate a product feature vector of product information.

[0058] Understandably, in the scenario of influencer live-streaming e-commerce, the merchants will introduce the products and answer users' questions about the products in the live-streaming video data or recorded live-streaming video data. The target text includes product information, etc.

[0059] In this embodiment of the disclosure, speech recognition is performed on video data to obtain target text; wherein, the target text includes product information; the target text is input into a trained deep language representation model to generate a product feature vector of the product information, thereby obtaining the product feature vector of the product information included in the target image.

[0060] In some other embodiments, processing the video data to generate a product feature vector of product information includes: extracting frames from the video data to obtain a target image, wherein the target image includes product information; performing speech recognition on the video data to obtain target text, wherein the target text includes product information; inputting the target image into a trained image processing model to generate an image feature vector; inputting the target text into a trained deep language representation model to generate a text feature vector; and concatenating the image feature vector and the text feature vector to generate a product feature vector of product information.

[0061] In this embodiment of the disclosure, when the target image is obtained by extracting frames from the video data and the target text is obtained by speech recognition from the video data, the target image in the video data is input into a residual network and a Vision Transformer network to generate an image feature vector; the target text in the video data is input into a deep language representation model to generate a text feature vector; the image feature vector and the text feature vector are concatenated to generate a product feature vector; wherein, placeholders are added to the corresponding positions of the target text in the input target image to jointly generate the product feature vector.

[0062] In this embodiment of the disclosure, there is a semantic relationship between the target image and the target text, and the target image and the target text exist in pairs.

[0063] In this embodiment of the disclosure, the target image and target text are processed differently to generate image feature vectors and text feature vectors. Furthermore, the image feature vectors and text feature vectors are concatenated by the concat() method to generate product feature vectors.

[0064] S23: Calculate the similarity between the product feature vector and multiple first-level category vectors in the target category prediction table of the first category level, and determine the multiple target first-level category vectors corresponding to the product feature vector.

[0065] In this embodiment of the disclosure, calculating the similarity between the product feature vector and multiple first-level category vectors in the target category prediction table of the first category level can be done by calculating the cosine distance between the product feature vector and multiple first-level category vectors in the target category prediction table of the first category level.

[0066] S24: Obtain multiple second-level category vectors from the target category prediction table at the second category level based on the target first-level category vector, and determine multiple target second-level category vectors corresponding to the product feature vector based on the similarity between the product feature vector and the determined multiple second-level category vectors.

[0067] S25: Continue in this manner until the similarity between the product feature vector and the determined multiple N-level category vectors is calculated, and a target N-level category vector corresponding to the product feature vector is determined; where N is the number of target category prediction tables.

[0068] S26: Determine the target level category of the product information based on the target N-level category vector.

[0069] In this embodiment of the disclosure, the target level category corresponding to the product information is determined based on the product feature vector and multiple target category prediction tables. The product feature vector and the target category prediction tables corresponding to different category levels are compared layer by layer. Based on the tree-like relationship between the category vectors in different target category prediction tables, the Beam Search method is used for each layer of the tree to sequentially determine multiple category vectors corresponding to different category levels of the product feature vector. Beam Search includes a parameter beamwidth, which means that at each time step, the sequence with the highest beamwidth score is retained, and then the next time step is used to continue generating with this beamwidth sequence. The beamwidth value is greater than or equal to 2.

[0070] In this embodiment, multiple target category vectors corresponding to different category levels of the product feature vector are determined sequentially, and the number of determined target category vectors is the beam width of the Beam Search. The Beam Search method used in this disclosure can determine the target category by combining the hierarchical tree relationship between category levels, thereby improving the accuracy of category classification.

[0071] For ease of understanding, the present disclosure provides the following embodiments:

[0072] The product category hierarchy comprises four levels: the first level includes 10 primary categories, the second level includes 50 secondary categories, the third level includes 500 tertiary categories, and the fourth level includes 1000 quaternary categories. The 10 primary categories in the first level generate a corresponding number of primary category vectors, which are stored in the target category prediction table for the first level. The 50 secondary categories in the second level generate a corresponding number of secondary category vectors, which are stored in the target category prediction table for the second level. Similarly, the 500 tertiary categories in the third level generate a corresponding number of tertiary category vectors, which are stored in the target category prediction table for the third level.

[0073] First, it is necessary to obtain the product feature vector and four target category prediction tables. The specific methods for obtaining these are described in the example above and will not be repeated here. In this embodiment, the Beam Search method is used, and the beam width of the Beam Search can be 2.

[0074] In this embodiment of the disclosure, the method for determining the target hierarchical category corresponding to product information based on product feature vectors and multiple target category prediction tables is as follows:

[0075] Step 1: Calculate the similarity between the product feature vector and the 10 first-level category vectors in the target category prediction table of the first category level, and determine the two target first-level category vectors corresponding to the product feature vector.

[0076] Step 2: In this embodiment of the present disclosure, multiple second-level category vectors in the target category prediction table of the second category level are obtained based on the two target first-level category vectors, and the two target second-level category vectors corresponding to the product feature vectors are determined based on the similarity between the product feature vector and the determined multiple second-level category vectors.

[0077] It is understandable that there is a tree-like relationship between the category vectors in different target category prediction tables. A tree-like relationship exists between first-level and second-level categories, meaning that multiple second-level category vectors can be determined based on two defined target first-level category vectors. The number of second-level category vectors corresponds to the number of second-level category vectors with which there is a tree-like relationship with the defined target first-level category.

[0078] Step 3: In this embodiment of the present disclosure, multiple third-level category vectors in the target category prediction table of the third category level are obtained based on the two target second-level category vectors, and the two target third-level category vectors corresponding to the product feature vectors are determined based on the similarity between the product feature vector and the determined multiple third-level category vectors.

[0079] Step 4: In this embodiment of the present disclosure, multiple fourth-level category vectors in the target category prediction table of the fourth category level are obtained based on the two target third-level category vectors, and one target fourth-level category vector corresponding to the product feature vector is determined based on the similarity between the product feature vector and the determined multiple fourth-level category vectors.

[0080] Understandably, in the example disclosed herein, the product category hierarchy includes four levels, and a category vector is ultimately determined when calculating the similarity of the last category level.

[0081] Step 4: Determine the target level category of the product information based on the target level 4 category vector.

[0082] It is understandable that the fourth-level category corresponding to the target fourth-level category vector, due to the tree-like relationship between category vectors in different target category prediction tables, can be used to sequentially determine the target third-level category vector, target second-level category vector, and target first-level category vector, thereby obtaining the target hierarchical category corresponding to the product information. In this embodiment of the disclosure, the Beam Search method is used, which can determine the target category by combining the hierarchical tree-like relationship between category levels, resulting in high accuracy in category classification.

[0083] like Figure 3 As shown, in some embodiments, the category classification method provided in this disclosure further includes:

[0084] S31: Randomly initialize multiple sample category prediction tables; wherein, each sample category prediction table includes multiple sample category vectors, each sample category prediction table corresponds to a category level, and the sample category vectors in different sample category prediction tables have a tree-like relationship.

[0085] In this embodiment of the disclosure, multiple sample category prediction tables are randomly initialized. Each sample category prediction table includes multiple sample category vectors. The length of each sample category vector is a specific length. Each sample category prediction table corresponds to a category level. The sample category vectors in different sample category prediction tables have a tree-like relationship.

[0086] S32: Obtain the training sample set; the training sample set includes multiple sample product information and the corresponding sample hierarchical categories.

[0087] The training sample set includes at least one set of sample product information labeled with sample hierarchical categories. The sample product information is in the form of sample image-text pairs, where each pair consists of a sample image and sample text, and these pairs exist in pairs with a semantic relationship. Alternatively, the sample product information may be in the form of sample images, and the training sample set includes at least one set of sample images labeled with sample hierarchical categories; or the sample product information may be in the form of sample text, and the training sample set includes at least one set of sample text labeled with sample hierarchical categories.

[0088] In this embodiment of the disclosure, the labeling of sample images and / or sample text with sample hierarchical categories can be done manually or in any other feasible manner, and this embodiment of the disclosure does not impose any specific restrictions on this.

[0089] S33: Process the sample product information to generate a sample product feature vector.

[0090] It is understood that the sample product information can be a sample image-text pair, a sample image, or a sample text. In this embodiment of the disclosure, the method for processing the sample product information to generate the sample product feature vector can be referred to the method for processing video data to generate the product feature vector of product information in the above example. When the data formats in the sample product information and the video data are the same, the same processing method is used to obtain the sample product feature vector of the sample product information, which will not be repeated here.

[0091] S34: Based on the sample product feature vector and multiple sample category prediction tables, determine the prediction target level category corresponding to the sample product information.

[0092] In this embodiment of the disclosure, the method for determining the predicted target category level of sample product information based on sample product feature vectors and multiple sample category prediction tables can be found in the description of the method for determining the target category level of product information based on product feature vectors and multiple target category prediction tables in the above example, and will not be repeated here.

[0093] S35: Calculate the sample loss value based on the predicted target level category and the sample level category.

[0094] It is understood that, in this embodiment of the disclosure, after determining the prediction target level category corresponding to the sample product information, the sample loss value is calculated based on the prediction target level category and the pre-labeled sample level category.

[0095] S36: Update the sample category prediction table based on the sample loss value and generate the target category prediction table.

[0096] In this embodiment, the process involves sequentially determining the target category for each sample product information based on its feature vector and labeled sample hierarchical categories in the training dataset, using the sample product feature vector and multiple sample category prediction tables. Then, a sample loss value is calculated based on the target category and the sample category. Finally, the sample category prediction tables are updated based on the sample loss values ​​to generate a target category prediction table. The target category prediction table obtained by performing the above process based on the previous set of sample feature vectors and labeled sample hierarchical categories is used as the next set of sample feature vectors and labeled sample hierarchical categories. Therefore, by using multiple sets of sample feature vectors and labeled sample hierarchical categories from the training dataset to generate a target category prediction table, the generated target category prediction table achieves high accuracy in classifying product information.

[0097] In some embodiments, when the sample product information is input into the image processing model to generate the sample product feature vector of the sample product information, the image processing model is updated according to the sample loss value to generate the trained image processing model.

[0098] Alternatively, when the sample product information is input into the deep language representation model to generate the sample product feature vector, the deep language representation model is updated according to the sample loss value to generate the trained deep language representation model.

[0099] Alternatively, when the sample product information is input into the image processing model and the deep language representation model to generate the sample product feature vector, the image processing model and the deep language representation model are updated according to the sample loss value to generate the trained image processing model and the trained deep language representation model.

[0100] In this embodiment of the disclosure, different vector generation models are used for sample product information of different formats. While updating the sample category prediction table according to the sample loss value, the different vector generation models used for sample product information of different formats are updated synchronously. This allows for the acquisition of trained vector generation models for sample product information of different formats, which can then be applied to the subsequent process of generating product feature vectors based on video data. This enables the accurate acquisition of the target hierarchical category corresponding to the product information in the video data.

[0101] It should be noted that the description of the above examples of the embodiments of this disclosure can be found in the relevant descriptions in the above examples, and will not be repeated here.

[0102] like Figure 4 As shown, in some embodiments, the category classification method provided in this disclosure further includes:

[0103] S4: Update the target category prediction table based on the addition and / or reduction of categories.

[0104] In this embodiment of the disclosure, when there are new or reduced categories, the target category prediction table is updated to obtain an updated target category prediction table. This allows the updated target category prediction table to be applied to determine the corresponding target level category based on the product feature vector, resulting in a more accurate determination of the target level category. Furthermore, it can be updated quickly when there are new or reduced categories, and the category classification method provided in this embodiment of the disclosure has strong scalability.

[0105] like Figure 5 As shown, in some embodiments, S41 includes:

[0106] S41: Randomly initialize multiple new sample category prediction tables; wherein, each new sample category prediction table includes multiple new sample category vectors, each new sample category prediction table corresponds to a category level, and the new sample category vectors in different new sample category prediction tables have a tree-like relationship.

[0107] In this embodiment of the disclosure, multiple new sample category prediction tables are randomly initialized. Each new sample category prediction table includes multiple new sample category vectors. The length of each new sample category vector is a specific length. Each new sample category prediction table corresponds to a category level. The new sample category vectors in different new sample category prediction tables have a tree-like relationship.

[0108] S42: Obtain the new training sample set for the new category; the new training sample set includes multiple new sample product information and the new sample hierarchical category corresponding to the new sample product information.

[0109] The newly added training sample set includes at least one set of newly added sample product information labeled with the new sample hierarchical category. The newly added sample product information is a newly added sample image-text pair, which includes both a newly added sample image and newly added sample text. The newly added sample image and newly added sample text exist in pairs and are semantically related. Alternatively, the newly added sample product information is a newly added sample image, and the newly added training sample set includes at least one set of newly added sample images labeled with the new sample hierarchical category; or, the newly added sample product information is newly added sample text, and the newly added training sample set includes at least one set of newly added sample text labeled with the new sample hierarchical category.

[0110] In this embodiment of the disclosure, the annotation of new sample images and / or new sample texts to add new sample hierarchical categories can be done manually or by any other feasible method. This embodiment of the disclosure does not impose any specific restrictions on this.

[0111] S43: Process the information of newly added sample products to generate a feature vector of newly added sample products.

[0112] It is understood that the newly added sample product information can be a newly added sample image-text pair, a newly added sample image, or a newly added sample text. In this embodiment of the disclosure, the method for processing the newly added sample product information to generate the newly added sample product feature vector can refer to the method for processing video data to generate product feature vectors in the above example. For cases where the data formats in the newly added sample product information and video data are the same, the same processing method is used to obtain the newly added sample product feature vector, which will not be elaborated here.

[0113] S44: Based on the feature vector of the newly added sample products and multiple prediction tables for the categories of newly added samples, determine the new prediction target category level corresponding to the information of the newly added sample products.

[0114] In this embodiment of the disclosure, the method for determining the new target level category corresponding to the new sample product information based on the new sample product feature vector and multiple new sample category prediction tables can be found in the relevant description of the method for determining the target level category corresponding to product information based on the product feature vector and multiple target category prediction tables in the above example, and will not be repeated here.

[0115] S45: Calculate the loss value of the new samples based on the newly added prediction target level category and the newly added sample level category.

[0116] It is understood that, in this embodiment of the disclosure, after determining the new prediction target category corresponding to the new sample product information, the new sample loss value is calculated based on the new prediction target category and the pre-labeled new sample category.

[0117] S46: Update the new sample category prediction table based on the loss value of the new sample, and generate the new target category prediction table.

[0118] In some embodiments, the newly added target category prediction table includes at least one newly added category vector. Updating the target category prediction table based on the newly added target category prediction table includes: updating the target category prediction table with the newly added category vector from the newly added target category prediction table, and updating the target category prediction table. In this embodiment, based on each set of newly added sample feature vectors and labeled newly added sample hierarchical categories in the newly added training dataset, the following steps are performed: determining the newly predicted target hierarchical category corresponding to the newly added sample product information based on the newly added sample product feature vectors and multiple newly added sample category prediction tables; calculating the newly added sample loss value based on the newly predicted target hierarchical category and the newly added sample hierarchical category; and updating the newly added sample category prediction table based on the newly added sample loss value to generate the newly added target category prediction table. The newly added target category prediction table obtained by performing the above process based on the previous set of newly added sample feature vectors and labeled newly added sample hierarchical categories is used as the newly added sample category prediction table for the next set of newly added sample feature vectors and labeled newly added sample hierarchical categories. Therefore, using multiple sets of new sample feature vectors and labeled new sample hierarchical categories from the newly added training dataset, a new target category prediction table is generated.

[0119] In this embodiment of the disclosure, after obtaining the new target category prediction table, the new target category prediction table can be directly added to the target category prediction table to obtain a new target category prediction table. This allows the new target category prediction table to be updated, so that the new target category prediction table can be used to classify the target level category of the new category. The update speed is fast and will not affect the category recognition accuracy.

[0120] In some embodiments, updating the target category prediction table based on the reduction of categories includes: deleting the category vectors corresponding to the reduced categories from the target category prediction table, and updating the target category prediction table.

[0121] In some embodiments, deleting the category vector corresponding to the reduced category from the target category prediction table includes: deleting the category vector corresponding to the reduced category in the first target category prediction table where the reduced category is located; and deleting the category vector that has a tree relationship with the first target category vector in at least one second target category prediction table after the category level corresponding to the first target category prediction table, based on the tree relationship between the category vectors in different target category prediction tables.

[0122] In an exemplary embodiment, the product category hierarchy includes three levels, which have a tree-like hierarchical relationship. The first-level category includes "Apparel". The second-level category corresponding to the first-level category "Apparel" includes "Men's Clothing" and "Women's Clothing". The third-level category corresponding to the second-level category "Men's Clothing" includes "Jeans", "Sweatshirts", and "Outerwear". The third-level category corresponding to the second-level category "Women's Clothing" includes "Jeans", "Dresses", and "Skirts". It should be understood that the above examples are for illustration only and are not intended to limit the embodiments of this disclosure.

[0123] Based on the above example, when reducing the category to "men's clothing", the category vector corresponding to "men's clothing" in the second-level target category prediction table is deleted. The third-level categories corresponding to the second-level category "men's clothing" include "jeans", "hoodies", and "outerwear". While deleting the category vector corresponding to the second-level category "men's clothing", the category vectors corresponding to the associated third-level categories "jeans", "hoodies", and "outerwear" can also be deleted.

[0124] Understandably, when the second-level category "Men's Clothing" is removed, the trained category classification model no longer needs to classify "Men's Clothing". Simply put, the trained category classification model only needs to classify the "Women's Clothing" category. The second-level categories corresponding to the third-level categories "Jeans", "Sweatshirts", and "Outerwear" are all "Men's Clothing". Since there's no need to classify the "Men's Clothing" category, the third-level categories corresponding to the removed second-level category "Men's Clothing" ("Jeans", "Sweatshirts", and "Outerwear") can be updated simultaneously, quickly reducing the number of categories and updating the target category prediction table.

[0125] In this embodiment of the disclosure, to reduce the number of categories, it is only necessary to delete the corresponding category vector in the target category prediction table. There is no need to re-acquire the target category prediction table, the update speed is fast, and it will not affect the category recognition accuracy, thus achieving the purpose of expanding the target category prediction table.

[0126] Figure 6 This is a block diagram illustrating a category classification device 1 according to an exemplary embodiment. (Refer to...) Figure 6 The device 1 includes: a data acquisition unit 11, a data processing unit 12, and a category determination unit 13.

[0127] The data acquisition unit 11 is used to acquire video data, which includes product information.

[0128] The data processing unit 12 is used to process video data and generate product feature vectors for product information.

[0129] The category determination unit 13 is used to determine the target level category corresponding to the product information based on the product feature vector and multiple target category prediction tables. Each target category prediction table includes multiple category vectors, each target category prediction table corresponds to a category level, and the category vectors in different target category prediction tables have a tree-like relationship.

[0130] like Figure 7 As shown, in some embodiments, the category determination unit 13 includes:

[0131] The first vector determination module 131 is used to calculate the similarity between the product feature vector and multiple first-level category vectors in the target category prediction table of the first category level, and to determine the multiple target first-level category vectors corresponding to the product feature vector.

[0132] The second vector determination module 132 is used to obtain multiple second-level category vectors in the target category prediction table of the second category level based on the target first-level category vector, and to determine multiple target second-level category vectors corresponding to the product feature vector based on the similarity between the product feature vector and the determined multiple second-level category vectors.

[0133] The third vector determination module 133 is used to calculate the similarity between the product feature vector and the determined multiple N-level category vectors, and determine a target N-level category vector corresponding to the product feature vector; where N is the number of target category prediction tables.

[0134] The first category determination module 134 is used to determine the target level category of the product information based on the target N-level category vector.

[0135] like Figure 8 As shown, in some embodiments, the data processing unit 12 includes:

[0136] The first processing module 121 is used to extract video frames from the video data to obtain a target image; wherein the target image includes product information.

[0137] The second processing module 122 is used to input the target image into the trained image processing model to generate a product feature vector of product information.

[0138] Please see Figure 8 In other embodiments, the data processing unit 12 includes:

[0139] The first processing module 121 is used to perform speech recognition on video data to obtain target text; wherein, the target text includes product information.

[0140] The second processing module 122 is used to input the target text into the trained deep language representation model to generate product feature vectors of product information.

[0141] like Figure 9 As shown, in some other embodiments, the data processing unit 12 includes:

[0142] The first processing module 121 is used to extract video frames from the video data to obtain a target image; wherein the target image includes product information.

[0143] The second processing module 122 is used to perform speech recognition on the video data and obtain target text; wherein, the target text includes product information.

[0144] The third processing module 123 is used to input the target image into the trained image processing model to generate image feature vectors.

[0145] The fourth processing module 124 is used to input the target text into the trained deep language representation model and generate text feature vectors;

[0146] The fifth processing module 125 is used to connect the image feature vector and the text feature vector to generate a product feature vector of product information.

[0147] Please see Figure 10 In some embodiments, the category classification device 1 further includes:

[0148] Initialization unit 14 is used to randomly initialize multiple sample category prediction tables; wherein, the sample category prediction table includes multiple sample category vectors, each sample category prediction table corresponds to a category level, and the sample category vectors in different sample category prediction tables have a tree-like relationship.

[0149] Dataset acquisition unit 15 is used to acquire training sample set; the training sample set includes multiple sample product information and sample hierarchical categories corresponding to the sample product information.

[0150] The sample processing unit 16 is used to process the sample product information and generate a sample product feature vector.

[0151] The prediction unit 17 is used to determine the prediction target level category corresponding to the sample product information based on the sample product feature vector and multiple sample category prediction tables.

[0152] The calculation unit 18 is used to calculate the sample loss value based on the prediction target level category and the sample level category.

[0153] The target generation unit 19 is used to update the sample category prediction table based on the sample loss value and generate the target category prediction table.

[0154] Please see Figure 11 In some embodiments, the category classification device 1 further includes:

[0155] The first model update unit 20 is used to update the image processing model according to the sample loss value when the sample product information is input into the image processing model to generate the sample product feature vector of the sample product information, and to generate the trained image processing model.

[0156] Alternatively, when the sample product information is input into the deep language representation model to generate the sample product feature vector, the deep language representation model is updated according to the sample loss value to generate the trained deep language representation model.

[0157] Alternatively, when the sample product information is input into the image processing model and the deep language representation model to generate the sample product feature vector, the image processing model and the deep language representation model are updated according to the sample loss value to generate the trained image processing model and the trained deep language representation model.

[0158] Please see again Figure 11 In some embodiments, the category classification device 1 further includes:

[0159] Update unit 21 is used to update the target category prediction table based on newly added and / or reduced categories.

[0160] Please see Figure 12 In some embodiments, the updating unit 21 includes:

[0161] A new initialization module 211 is added to randomly initialize multiple new sample category prediction tables. Each new sample category prediction table includes multiple new sample category vectors. Each new sample category prediction table corresponds to a category level, and the new sample category vectors in different new sample category prediction tables have a tree-like relationship.

[0162] The newly added dataset acquisition module 212 is used to acquire the newly added training sample set for the newly added category; the newly added training sample set includes information on multiple newly added sample products and the corresponding newly added sample hierarchical categories.

[0163] A new sample processing module 213 is added to process the information of newly added sample products and generate a feature vector of the newly added sample products.

[0164] A new prediction module 214 is added, which is used to determine the new prediction target level category corresponding to the new sample product information based on the feature vector of the new sample product and multiple new sample category prediction tables.

[0165] A new calculation module 215 is added to calculate the loss value of new samples based on the new prediction target level category and the new sample level category.

[0166] A new target generation module 216 is added, which is used to update the new sample category prediction table based on the loss value of the new sample and generate a new target category prediction table.

[0167] Please see again Figure 12 In some embodiments, the updating unit 21 includes:

[0168] A new target update module 217 is added to update the target category prediction table by adding new category vectors from the new target category prediction table; wherein, the new target category prediction table includes at least one new category vector.

[0169] Please see again Figure 12 In some embodiments, the updating unit 21 includes:

[0170] The first deletion module 218 is used to delete the category vector corresponding to the reduced category from the target category prediction table and update the target category prediction table.

[0171] In some embodiments, the first deletion module 218 is specifically used to delete the category vector corresponding to the reduced category in the first target category prediction table where the reduced category is located; and according to the tree relationship between the category vectors in different target category prediction tables, to delete the category vectors that have a tree relationship with the first target category vector in at least one second target category prediction table after the category level corresponding to the first target category prediction table.

[0172] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0173] By implementing the category classification method provided in this embodiment, the data acquisition unit 11 acquires video data, which includes product information. The data processing unit 12 processes the video data to generate product feature vectors for the product information. The category determination unit 13 determines the target category level corresponding to the product information based on the product feature vectors and multiple target category prediction tables. Each target category prediction table includes multiple category vectors, each target category prediction table corresponds to a category level, and the category vectors in different target category prediction tables have a tree-like relationship. This improves the accuracy of category classification.

[0174] Figure 13 This is a block diagram illustrating an electronic device 100 for a category classification method or a category classification model training method, according to an exemplary embodiment.

[0175] For example, electronic device 100 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.

[0176] like Figure 13 As shown, the electronic device 100 may include one or more of the following components: a processing component 101, a memory 102, a power supply component 103, a multimedia component 104, an audio component 105, an input / output (I / O) interface 106, a sensor component 107, and a communication component 108.

[0177] Processing component 101 typically controls the overall operation of electronic device 100, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 101 may include one or more processors 1011 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 101 may include one or more modules to facilitate interaction between processing component 101 and other components. For example, processing component 101 may include a multimedia module to facilitate interaction between multimedia component 104 and processing component 101.

[0178] Memory 102 is configured to store various types of data to support the operation of electronic device 100. Examples of such data include instructions for any application or method operating on electronic device 100, contact data, phonebook data, messages, pictures, videos, etc. Memory 102 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as SRAM (Static Random-Access Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM (Erasable Programmable Read-Only Memory), PROM (Programmable Read-Only Memory), ROM (Read-Only Memory), magnetic storage, flash memory, magnetic disk, or optical disk.

[0179] Power supply component 103 provides power to various components of electronic device 100. Power supply component 103 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 100.

[0180] Multimedia component 104 includes a touch display screen that provides an output interface between the electronic device 100 and the user. In some embodiments, the touch display screen may include an LCD (Liquid Crystal Display) and a TP (Touch Panel). The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 104 includes a front-facing camera and / or a rear-facing camera. When the electronic device 100 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0181] Audio component 105 is configured to output and / or input audio signals. For example, audio component 105 includes a MIC (microphone) configured to receive external audio signals when electronic device 100 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 102 or transmitted via communication component 108. In some embodiments, audio component 105 also includes a speaker for outputting audio signals.

[0182] I / O interface 2112 provides an interface between processing component 101 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, start buttons, and lock buttons.

[0183] Sensor assembly 107 includes one or more sensors for providing state assessments of various aspects of electronic device 100. For example, sensor assembly 107 may detect the on / off state of electronic device 100, the relative positioning of components such as the display and keypad of electronic device 100, changes in position of electronic device 100 or a component of electronic device 100, the presence or absence of user contact with electronic device 100, orientation or acceleration / deceleration of electronic device 100, and temperature changes of electronic device 100. Sensor assembly 107 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 107 may also include an optical sensor, such as a CMOS (Complementary Metal Oxide Semiconductor) or CCD (Charge-coupled Device) image sensor, for use in imaging applications. In some embodiments, sensor assembly 107 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.

[0184] Communication component 108 is configured to facilitate wired or wireless communication between electronic device 100 and other devices. Electronic device 100 can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 108 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 108 also includes an NFC (Near Field Communication) module to facilitate short-range communication. For example, the NFC module may be implemented based on RFID (Radio Frequency Identification) technology, IrDA (Infrared Data Association) technology, UWB (Ultra Wide Band) technology, BT (Bluetooth) technology, and other technologies.

[0185] In an exemplary embodiment, the electronic device 100 may be implemented by one or more ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), digital signal processing devices (DSPDs), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), controllers, microcontrollers, microprocessors, or other electronic components to execute the above-described category classification method. It should be noted that the implementation process and technical principles of the electronic device in this embodiment are explained in the foregoing description of the category classification method of this disclosure, and will not be repeated here.

[0186] The electronic device provided in this disclosure can perform the category classification method as described in some of the above embodiments, and its beneficial effects are the same as those of the above category classification method, which will not be repeated here.

[0187] To implement the above embodiments, this disclosure also proposes a storage medium.

[0188] When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the category classification method as described above. For example, the storage medium may be ROM (Read Only Memory Image), RAM (Random Access Memory), CD-ROM (Compact Disc Read-Only Memory), magnetic tape, floppy disk, and optical data storage device, etc.

[0189] To implement the above embodiments, this disclosure also provides a computer program product that, when executed by the processor of an electronic device, enables the electronic device to perform the category classification method as described above.

[0190] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0191] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A category classification method, characterized in that, The method includes: Acquire video data; wherein the video data includes product information; The video data is subjected to frame extraction to obtain a target image, and the video data is subjected to speech recognition to obtain target text; wherein the target image and the target text include the product information, the target image and the target text have a semantic relationship, and the target image and the target text exist in pairs; The target image is input into a trained image processing model to generate an image feature vector. The target text is input into a trained deep language representation model to generate a text feature vector. The image feature vector and the text feature vector are concatenated to generate the product feature vector of the product information; Calculate the similarity between the product feature vector and multiple first-level category vectors in the target category prediction table of the first category level. Based on the preset bundle width, use the bundle search method to select the K target first-level category vectors with the highest similarity to the product feature vector from the multiple first-level category vectors, where K is an integer greater than 1 and equal to the bundle width. Based on the tree-like relationship between the K target first-level category vectors and the category vectors in different target category prediction tables, determine and obtain multiple second-level category vectors corresponding to the target category prediction table at the second category level; based on the similarity between the product feature vector and the multiple second-level category vectors, and also based on the bundle width, use the bundle search method to select K target second-level category vectors from the multiple second-level category vectors; This process continues until the similarity between the product feature vector and the determined multiple N-level category vectors is calculated, thereby determining a target N-level category vector corresponding to the product feature vector; where N is the number of target category prediction tables. Based on the target N-level category vector, the target hierarchical category of the product information is determined; wherein, each target category prediction table includes multiple category vectors, each target category prediction table corresponds to a category level, and the category vectors in different target category prediction tables have a tree-like relationship.

2. The method according to claim 1, characterized in that, The method further includes: Multiple sample category prediction tables are randomly initialized; wherein, each sample category prediction table includes multiple sample category vectors, each sample category prediction table corresponds to a category level, and the sample category vectors in different sample category prediction tables have a tree-like relationship; Obtain a training sample set; the training sample set includes multiple sample product information and the sample hierarchical categories corresponding to the sample product information; The sample product information is processed to generate a sample product feature vector of the sample product information; Based on the sample product feature vector and multiple sample category prediction tables, the prediction target level category corresponding to the sample product information is determined; Calculate the sample loss value based on the predicted target hierarchical category and the sample hierarchical category; The sample category prediction table is updated based on the sample loss value to generate the target category prediction table.

3. The method according to claim 2, characterized in that, The method further includes: When the sample product information is input into the image processing model to generate the sample product feature vector of the sample product information, the image processing model is updated according to the sample loss value to generate a trained image processing model. Alternatively, when the sample product information is input into a deep language representation model to generate a sample product feature vector of the sample product information, the deep language representation model is updated according to the sample loss value to generate a trained deep language representation model. Alternatively, when the sample product information is input into the image processing model and the deep language representation model to generate the sample product feature vector of the sample product information, the image processing model and the deep language representation model are updated according to the sample loss value to generate the trained image processing model and the trained deep language representation model.

4. The method according to claim 1, characterized in that, The method further includes: The target category prediction table is updated based on the addition and / or reduction of categories.

5. The method according to claim 4, characterized in that, The step of updating the target category prediction table based on the newly added categories includes: Multiple new sample category prediction tables are randomly initialized; wherein, each new sample category prediction table includes multiple new sample category vectors, each new sample category prediction table corresponds to a category level, and the new sample category vectors in different new sample category prediction tables have a tree-like relationship. Obtain the newly added training sample set for the newly added category; the newly added training sample set includes multiple newly added sample product information and the newly added sample hierarchical category corresponding to the newly added sample product information; The newly added sample product information is processed to generate a new sample product feature vector. Based on the newly added sample product feature vector and multiple newly added sample category prediction tables, determine the newly added prediction target level category corresponding to the newly added sample product information; Calculate the loss value of the new sample based on the newly added prediction target hierarchical category and the newly added sample hierarchical category; The new sample category prediction table is updated based on the new sample loss value to generate the new target category prediction table.

6. The method according to claim 5, characterized in that, The newly added target category prediction table includes at least one newly added category vector, and the method further includes: The newly added category vector in the newly added target category prediction table is updated to the target category prediction table, thereby updating the target category prediction table.

7. The method according to any one of claims 4 to 6, characterized in that, The step of updating the target category prediction table based on the reduction of categories includes: The category vector corresponding to the reduced category is deleted from the target category prediction table, and the target category prediction table is updated.

8. The method according to claim 7, characterized in that, The step of deleting the category vector corresponding to the reduced category from the target category prediction table includes: Delete the category vector corresponding to the reduced category from the first target category prediction table where the reduced category is located; and Based on the tree-like relationship between the category vectors in different target category prediction tables, the category vectors that have a tree-like relationship with the first target category vector in at least one second target category prediction table after the category level corresponding to the first target category prediction table are deleted.

9. A category classification device, characterized in that, The device includes: A data acquisition unit is used to acquire video data; wherein the video data includes product information; A data processing unit is configured to extract target images from the video data by performing video frame extraction and to obtain target text from the video data by performing speech recognition; wherein the target image and the target text include product information, the target image and the target text have a semantic relationship, and the target image and the target text exist in pairs; the target image is input into a trained image processing model to generate an image feature vector; the target text is input into a trained deep language representation model to generate a text feature vector; and the image feature vector and the text feature vector are concatenated to generate a product feature vector for the product information; The category determination unit is used to determine the target level category corresponding to the product information based on the product feature vector and multiple target category prediction tables; wherein, each target category prediction table includes multiple category vectors, each target category prediction table corresponds to a category level, and the category vectors in different target category prediction tables have a tree-like relationship. The category determination unit includes: The first vector determination module is used to calculate the similarity between the product feature vector and multiple first-level category vectors in the target category prediction table of the first category level. Based on the preset bundle width, the bundle search method is used to select the K target first-level category vectors with the highest similarity to the product feature vector from the multiple first-level category vectors, where K is an integer greater than 1 and equal to the bundle width. The second vector determination module is used to determine and obtain multiple second-level category vectors corresponding to the target category prediction table at the second category level based on the tree-like relationship between the K target first-level category vectors and the category vectors in different target category prediction tables; and to select K target second-level category vectors from the multiple second-level category vectors based on the similarity between the product feature vector and the multiple second-level category vectors, and also based on the bundle width, using a bundle search method. The third vector determination module is used to calculate the similarity between the product feature vector and the determined multiple N-level category vectors, and determine a target N-level category vector corresponding to the product feature vector; where N is the number of target category prediction tables. The first category determination module is used to determine the target hierarchical category of the product information based on the target N-level category vector.

10. The apparatus according to claim 9, characterized in that, The device further includes: An initialization unit is used to randomly initialize multiple sample category prediction tables; wherein, each sample category prediction table includes multiple sample category vectors, each sample category prediction table corresponds to a category level, and the sample category vectors in different sample category prediction tables have a tree-like relationship. A dataset acquisition unit is used to acquire a training sample set; the training sample set includes multiple sample product information and the sample hierarchical categories corresponding to the sample product information. A sample processing unit is used to process the sample product information and generate a sample product feature vector of the sample product information. The prediction unit is used to determine the prediction target level category corresponding to the sample product information based on the sample product feature vector and multiple sample category prediction tables. The calculation unit is used to calculate the sample loss value based on the predicted target hierarchical category and the sample hierarchical category; The target generation unit is used to update the sample category prediction table based on the sample loss value and generate the target category prediction table.

11. The apparatus according to claim 10, characterized in that, The device further includes: The first model update unit is used to update the image processing model according to the sample loss value when the sample product information is input into the image processing model to generate the sample product feature vector of the sample product information, thereby generating a trained image processing model. Alternatively, when the sample product information is input into a deep language representation model to generate a sample product feature vector of the sample product information, the deep language representation model is updated according to the sample loss value to generate a trained deep language representation model. Alternatively, when the sample product information is input into the image processing model and the deep language representation model to generate the sample product feature vector of the sample product information, the image processing model and the deep language representation model are updated according to the sample loss value to generate the trained image processing model and the trained deep language representation model.

12. The apparatus according to claim 9, characterized in that, The device further includes: The update unit is used to update the target category prediction table based on the addition and / or reduction of categories.

13. The apparatus according to claim 12, characterized in that, The update unit includes: An initialization module is added to randomly initialize multiple new sample category prediction tables; wherein, each new sample category prediction table includes multiple new sample category vectors, each new sample category prediction table corresponds to a category level, and the new sample category vectors in different new sample category prediction tables have a tree-like relationship. A new dataset acquisition module is used to acquire the new training sample set for the new category; the new training sample set includes multiple new sample product information and the new sample hierarchical category corresponding to the new sample product information. A new sample processing module is added to process the new sample product information and generate a new sample product feature vector of the new sample product information; A new prediction module is added to determine the new prediction target level category corresponding to the new sample product information based on the feature vector of the new sample product and multiple new sample category prediction tables. A new calculation module is added to calculate the loss value of the new samples based on the new prediction target level category and the new sample level category; A new target generation module is added, which is used to update the new sample category prediction table based on the new sample loss value and generate a new target category prediction table.

14. The apparatus according to claim 13, characterized in that, The update unit includes: A new target update module is added, which is used to update the target category prediction table by adding the new category vector in the new target category prediction table; wherein, the new target category prediction table includes at least one new category vector.

15. The apparatus according to any one of claims 12 to 14, characterized in that, The update unit includes: The first deletion module is used to delete the category vector corresponding to the reduced category from the target category prediction table and update the target category prediction table.

16. The apparatus according to claim 15, characterized in that, The first deletion module is specifically used to delete the category vector corresponding to the reduced category in the first target category prediction table where the reduced category is located; and according to the tree relationship between the category vectors in different target category prediction tables, to delete the category vectors that have a tree relationship with the first target category vector in at least one second target category prediction table after the category level corresponding to the first target category prediction table.

17. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the category classification method as described in any one of claims 1 to 8.

18. A storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the category classification method as described in any one of claims 1 to 8.

19. A computer program product comprising a computer program that, when executed by a processor, implements the category classification method according to any one of claims 1 to 8.