Method and apparatus, device, medium, and product for extracting keywords from product titles

Through the text classification model, the candidate keywords in the product title are classified, which solves the problem of redundant information in the product title and improves the product information matching efficiency of e-commerce platforms.

CN114818674BActive Publication Date: 2025-07-25BUSINESS LINE COMMERCIAL PTE LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210501438.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-09
Publication Date
2025-07-25
Estimated Expiration
2042-05-09

AI Technical Summary

Technical Problem

The prior art cannot effectively extract keywords from product titles of e-commerce websites, resulting in excessive redundant information and affecting the efficiency of product search, advertising and recommendations.

Method used

The text classification model is used to classify the candidate keywords and the title text sentence pairs in the product title, and the sentence pairs with relevant categories as target related categories are selected, and the candidate keywords are determined as target keywords.

Benefits of technology

Through semantic recognition capabilities, redundant information is filtered, target keywords with high information value are extracted, and product information matching efficiency of e-commerce platforms is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114818674B_ABST
    Figure CN114818674B_ABST
Patent Text Reader

Abstract

This application relates to a method and apparatus, device, medium, and product for extracting keyword phrases from product titles in the field of e-commerce information technology. The method includes: obtaining the title text of a product; extracting candidate keyword phrases belonging to product words and attribute words from the title text; forming sentence pairs by combining each candidate keyword phrase with the title text, and using a text classification model that has been trained to a convergent state to classify each sentence pair respectively, so as to determine the relevant category representing the degree of relevance between the candidate keyword phrase in each sentence pair and the title text; screening out the sentence pairs whose relevant category is the target relevant category, and using the candidate keyword phrases therein as the target keyword phrases of the title text. This application utilizes the semantic recognition ability obtained through training of the text classification model to filter redundant information in the title text, identify target keyword phrases with high information value, and improve the efficiency of matching product information on e-commerce platforms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of e-commerce information processing, and in particular, to a method for extracting keywords from a product title, and a corresponding device, computer device, computer-readable storage medium, and computer program product. Background Art

[0002] For products sold on e-commerce websites, a lot of descriptive words are usually piled up in their titles to increase the traffic of SEO (Search Engine Optimization), resulting in very long product titles with certain redundant information, or even information unrelated to the products. In algorithm models for product search, advertising, and recommendation, the product title is important input information, and it is necessary to filter out the redundancy and noise therein and extract key summary information.

[0003] There are some solutions in natural language processing technology for extracting keywords from a piece of text or a long text, but these technical solutions do not consider the particularity of product titles on e-commerce websites, and thus cannot be directly used to process product titles.

[0004] A product title is generally a sentence piled up with many words and without a complete grammatical structure, which is quite different from the texts processed by the existing keyword extraction-related technologies. Therefore, how to extract keywords from a product title remains an urgent problem to be solved. Summary of the Invention

[0005] The primary objective of this application is to solve at least one of the above problems, and to provide a method for extracting keywords from a product title, and a corresponding device, computer device, computer-readable storage medium, and computer program product.

[0006] To achieve the various objectives of this application, the following technical solutions are adopted:

[0007] A method for extracting keywords from a product title provided to meet one of the objectives of this application includes the following steps:

[0008] Obtain the title text of the product;

[0009] Extract candidate keywords belonging to product words and attribute words from the title text;

[0010] Form sentence pairs by combining each candidate keyword with the title text, and use a text classification model that has been trained to a convergent state to classify each sentence pair respectively, and determine the relevant category representing the relevance between the candidate keyword in each sentence pair and the title text;

[0011] Filter out the sentence pairs whose relevant categories are the target relevant categories, and use the candidate keywords therein as the target keywords of the title text.

[0012] In a deepened partial embodiment, extracting candidate keywords from the title text includes the following steps:

[0013] Match the title text with a preset product word library to obtain the product words in the title text;

[0014] Match the title text with a preset attribute word library to obtain the attribute words in the title text;

[0015] Determine the product words and attribute words as the candidate keywords of the title text.

[0016] In an extended partial embodiment, before the step of determining the product words and attribute words as the candidate keywords of the title text, the following steps are included:

[0017] Filter out the product words or attribute words with semantic similarity lower than a preset threshold according to the semantic similarity between the title text and its product words or attribute words.

[0018] In an extended partial embodiment, before the step of using a text classification model that has been trained to a convergent state to classify each sentence pair formed by each candidate keyword and the title text, the following steps are included:

[0019] Use the training samples in a preset dataset to perform iterative training on the text classification model and train it to a convergent state. The training samples include the title text and a single candidate keyword included in the title text.

[0020] In an extended partial embodiment, after the step of using the candidate keywords therein as the target keywords of the title text, the following steps are included:

[0021] Determine the word frequency feature of each target keyword of the title text according to the statistical word frequency of each target keyword of the title text in a preset title library;

[0022] Determine the position feature of each target keyword of the title text according to the position of each target keyword of the title text in the title text;

[0023] Quantify the word frequency feature and position feature of each target keyword of the title text to determine the information score of the target keyword;

[0024] Select the combined text of the product words and the attribute words as the title abstract of the commodity according to the information score.

[0025] In some enhanced embodiments, the word frequency feature of each target keyword in the title text is determined according to its statistical word frequency in a preset title library, including the following steps:

[0026] The first word frequency feature of each target keyword is determined according to its statistical word frequency in the first title library, where the first title library is composed of the title texts of products of the same category as the product;

[0027] The second word frequency feature of each target keyword is determined according to its statistical word frequency in the second title library, where the second title library is composed of the title texts of products in the same online store as the product.

[0028] In some enhanced embodiments, the position feature of each target keyword in the title text is determined according to its position in the title text, including the following steps:

[0029] The absolute position feature of each target keyword is determined according to its absolute position in the title text;

[0030] The relative position feature of each target keyword belonging to an attribute word is determined according to its relative position in the title text relative to its closest product word;

[0031] For each target keyword belonging to a product word, its relative position feature is determined with a standard value.

[0032] A device for extracting keywords from a product title provided to meet one of the purposes of the present application includes a title acquisition module, a term extraction module, a term classification module, and a target determination module, where: the title acquisition module is used to acquire the title text of a product; the term extraction module is used to extract candidate keywords belonging to product words and attribute words from the title text; the term classification module is used to form sentence pairs with each candidate keyword and the title text, and use a text classification model that has been trained to a convergent state to classify each sentence pair respectively, and determine the relevant category representing the relevance between the candidate keyword in each sentence pair and the title text; the target determination module is used to screen out the sentence pairs whose relevant category is the target relevant category, and use the candidate keywords therein as the target keywords of the title text.

[0033] In some enhanced embodiments, the term extraction module includes: a product word extraction unit for matching the title text with a preset product word library to obtain the product words in the title text; an attribute word extraction unit for matching the title text with a preset attribute word library to obtain the attribute words in the title text; and a candidate set determination unit for determining the product words and attribute words as the candidate keywords of the title text.

[0034] In an extended partial embodiment, prior to the candidate set determination unit, it includes: a similarity filtering unit, configured to filter product words or attribute words with a semantic similarity lower than a preset threshold according to the semantic similarity between the title text and its product words or attribute words.

[0035] In an extended partial embodiment, prior to the entry classification module, it includes: a model training module, configured to perform iterative training on a text classification model using training samples in a preset dataset until it converges. The training samples include title texts and individual candidate keywords included in the title texts.

[0036] In an extended partial embodiment, after the target determination unit, it includes: a word frequency feature determination unit, configured to determine the word frequency feature of each target keyword in the title text according to the statistical word frequency in a preset title library; a position feature determination unit, configured to determine the position feature of each target keyword in the title text according to its position in the title text; an information score determination unit, configured to quantify the word frequency feature and position feature of each target keyword in the title text to determine the information score of the target keyword; a title abstract selection unit, configured to select the combined text of the product words and the attribute words as the title abstract of the commodity according to the information score.

[0037] In a deepened partial embodiment, the word frequency feature determination unit includes: a first word frequency feature unit, configured to determine the first word frequency feature of each target keyword according to the statistical word frequency in a first title library, where the first title library is a title library composed of title texts of commodities of the same category as the commodity; a second word frequency feature unit, configured to determine the second word frequency feature of each target keyword according to the statistical word frequency in a second title library, where the second title library is a title library composed of title texts of commodities in the same online store as the commodity.

[0038] In a deepened partial embodiment, the position feature determination unit includes: an absolute position feature unit, configured to determine the absolute position feature of each target keyword according to its absolute position in the title text; a relative position unit for attribute words, configured to determine the relative position feature of each target keyword belonging to an attribute word according to its relative position relative to the closest product word in the title text; a relative position unit for product words, configured to determine the relative position feature of each target keyword belonging to a product word with a standard value.

[0039] A computer device provided to meet one of the purposes of the present application includes a central processing unit and a memory. The central processing unit is configured to call and run a computer program stored in the memory to execute the steps of the commodity title keyword extraction method described in the present application.

[0040] A computer-readable storage medium provided for another purpose of the present application stores a computer program implemented according to the above-mentioned method for extracting product title keywords in the form of computer-readable instructions. When the computer program is called and run by a computer, it executes the steps included in the method.

[0041] A computer program product provided for another purpose of the present application includes a computer program / instructions. When the computer program / instructions are executed by a processor, they implement the steps of the method described in any embodiment of the present application.

[0042] Compared with the prior art, the advantages of the present application include:

[0043] According to the characteristic that the product title is composed of multiple piled-up words in the present application, sentence pairs are formed by using the title text of the product title and the candidate keywords in the title text. Then, a text classification model with the ability to classify the degree of semantic relevance learned in advance is used to classify and judge each sentence pair to determine the relevant category of the candidate keyword and the title text. Finally, the candidate keywords in the title text are selected as target keywords according to the preset target relevant category. Thus, by using the semantic recognition ability obtained through model training, redundant information in the title text is filtered, and target keywords with high information value are hit, which is applicable to providing basic materials required for product search, product advertising, and product recommendation in an e-commerce platform, thereby improving the product information matching efficiency of the entire e-commerce platform. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The above and / or additional aspects and advantages of the present application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0045] Figure 1 is a schematic flowchart of a typical embodiment of the method for extracting product title keywords of the present application;

[0046] Figure 2 is a schematic flowchart of the process of extracting candidate keywords in an embodiment of the present application;

[0047] Figure 3 is a schematic flowchart of the process of determining the title abstract of the title text in an embodiment of the present application;

[0048] Figure 4 is a schematic flowchart of the process of determining multiple word frequency features in an embodiment of the present application;

[0049] Figure 5 is a schematic flowchart of the process of determining multiple position features in an embodiment of the present application;

[0050] Figure 6Schematic flowchart of the process for screening title abstracts in an embodiment of the present application;

[0051] Figure 7 Schematic flowchart of the process for screening title abstracts in another embodiment of the present application;

[0052] Figure 8 Principle block diagram of the device for extracting keywords from product titles in the present application;

[0053] Figure 9 Schematic diagram of the structure of a computer device adopted in the present application. Detailed implementation manners

[0054] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals indicate the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be construed as a limitation to the present application.

[0055] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0056] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.

[0057] Those skilled in the art of the present technology can understand that the "client", "terminal", and "terminal device" used herein include both devices with wireless signal receivers that only have the ability to receive and do not have the ability to transmit, and devices with receiving and transmitting hardware that have the receiving and transmitting hardware capable of two-way communication on a two-way communication link. Such devices may include: cellular or other communication devices such as personal computers, tablet computers, etc., which have a single-line display or a multi-line display or a cellular or other communication device without a multi-line display; PCS (Personal Communications Service), which can combine voice, data processing, fax, and / or data communication capabilities; PDA (Personal Digital Assistant), which may include a radio frequency receiver, pager, Internet / intranet access, web browser, notepad, calendar, and / or GPS (Global Positioning System) receiver; conventional laptop and / or palm computers or other devices, which are conventional laptop and / or palm computers or other devices with and / or including a radio frequency receiver. The "client", "terminal", and "terminal device" used herein can be portable, transportable, installed in a vehicle (air, sea, and / or land), or suitable for and / or configured to run locally, and / or run in a distributed manner at any other location on the earth and / or in space. The "client", "terminal", and "terminal device" used herein can also be a communication terminal, an Internet access terminal, a music / video playback terminal, for example, it can be a PDA, MID (Mobile Internet Device), and / or a mobile phone with music / video playback function, or it can also be a smart TV, a set-top box, etc.

[0058] The hardware referred to by names such as "server", "client", and "service node" in this application is essentially an electronic device with the equivalent capabilities of a personal computer, and is a hardware device with the necessary components disclosed by the von Neumann principle, including a central processing unit (including an arithmetic unit and a controller), a memory, an input device, and an output device. The computer program is stored in its memory, and the central processing unit loads the program stored in the external memory into the internal memory for execution, executes the instructions in the program, and interacts with the input / output devices to complete specific functions.

[0059] It should be noted that the concept of "server" in this application can similarly be extended to the case applicable to a server cluster. According to the network deployment principle understood by those skilled in the art, the servers should be logically divided. Physically, these servers can either be independent of each other but can be called through an interface, or be integrated into a physical computer or a set of computer clusters. Those skilled in the art should understand this flexibility and should not be restricted by this when implementing the network deployment method of this application.

[0060] One or several technical features of this application, unless explicitly specified, can either be deployed on the server and accessed by the client remotely invoking the online service interface provided by the server, or directly deployed and run on the client for access.

[0061] The neural network models cited or possibly cited in this application, unless explicitly specified, can either be deployed on a remote server and remotely invoked by the client, or be directly invoked on the client with sufficient device capabilities. In some embodiments, when it runs on the client, its corresponding intelligence can be obtained through transfer learning to reduce the requirements for the client's hardware operation resources and avoid excessive occupation of the client's hardware operation resources.

[0062] All kinds of data involved in this application, unless explicitly specified, can either be remotely stored on the server or stored on the local terminal device, as long as it is suitable for being invoked by the technical solution of this application.

[0063] Those skilled in the art should be aware that although the various methods of this application are described based on the same concept and thus show commonality with each other, unless otherwise specified, these methods can be executed independently. Similarly, for the various embodiments disclosed in this application, they are all proposed based on the same inventive concept. Therefore, for the same expressed concepts, as well as concepts that are only appropriately transformed for convenience although the concept expressions are different, they should be equivalently understood.

[0064] For the various embodiments to be disclosed in this application, unless explicitly pointed out that there is a mutually exclusive relationship between them, the relevant technical features involved in each embodiment can be cross-combined to flexibly construct new embodiments, as long as this combination does not deviate from the creative spirit of this application and can meet the requirements in the prior art or solve certain deficiencies in the prior art. Those skilled in the art should be aware of this flexibility.

[0065] A method for extracting keywords from a product title of the present application can be programmed as a computer program product and implemented by running on a client or a server. For example, in the application scenario of the e-commerce platform of the present application, it is generally deployed on the server for implementation. Thus, by accessing the interface opened after the computer program product runs, human-computer interaction can be performed with the process of the computer program product through a graphical user interface to execute the method.

[0066] Please refer to Figure 1 , in a typical embodiment of the method for extracting keywords from a product title of the present application, the following steps are included:

[0067] Step S1100: Obtain the title text of the product:

[0068] Adapting to the differences in specific application scenarios, the title text of the product can be obtained from corresponding channels in each application scenario. For example:

[0069] In an exemplary application scenario, when a merchant user of an online store on an e-commerce platform publishes a product promotion activity, they need to fill in the copywriting of the product-related information. At this time, the product summary refined from the product title can be directly used as the writing material for the promotion copy.

[0070] In an exemplary application scenario, a merchant user of an online store on an e-commerce platform needs to publish the product information of the products they have launched, which includes the title text corresponding to the product title. The title text can be collected to generate a title summary for the merchant user to call.

[0071] In another exemplary application scenario, an e-commerce platform needs to determine the products that match the query information provided by consumer users, so as to implement functions such as product search, product advertising, and product recommendation. For this purpose, corresponding title summaries can be generated for the product titles in the product information of each product in the product database of the online store on the e-commerce platform first, and by matching the query information with the title summaries, some products can be determined to be recommended to consumer users.

[0072] In yet another exemplary application scenario, it is necessary to push the summary information of a certain product, including the title information of the product, to relevant users, including consumer users or merchant users. Therefore, a corresponding summary text can be generated based on the title file obtained from the product information of the product and encapsulated in the summary information and pushed to the user.

[0073] And so on. Adapting to the differences in specific application scenarios of the electronic product platform, the above-mentioned title text can be obtained from multiple channels.

[0074] Step S1200: Extract candidate keywords belonging to product words and attribute words from the title text:

[0075] Generally, the title text contains multiple entries. The main components of the entries include product words with noun attributes and attribute words with adjective or adverb attributes. The product words are mainly used to indicate the product name of the commodity or its substitute, and the attribute words are mainly used to describe information on certain aspects of the commodity's characteristics, functions, effects, etc. Various methods can be used to extract the corresponding entries of the product words and attribute words from the title text, and these entries constitute the candidate keywords in the title text.

[0076] For example, for clothing commodities, the product words include but are not limited to: skirt, dress, women's skirt, long dress, and the attribute words include but are not limited to: long sleeves, round neck, pure cotton, casual.

[0077] The method of extracting candidate keywords from the title text is relatively flexible. For example, based on rule matching or semantic matching, etc., the corresponding product words and attribute words can be extracted from the title text and temporarily stored in the corresponding product word list and attribute word list of the title text for subsequent calls. In addition, in some alternative embodiments, the title text can be segmented, and based on each segment, it is determined whether it belongs to a product word or an attribute word through different matching methods. When segmenting the title text, word segmentation tools such as Jieba, SnowNLP, and PkuSeg can be used.

[0078] Step S1300: Combine each candidate keyword with the title text to form a sentence pair, and use a text classification model that has been trained to a convergent state to classify each sentence pair respectively, and determine the relevant category representing the relevance between the candidate keyword in each sentence pair and the title text:

[0079] Each candidate keyword can form a corresponding sentence pair with the title text, and with the help of a prepared text classification model, the relevance between the candidate keyword in the sentence pair and the title text is judged based on this sentence pair.

[0080] For the convenience of model processing, each sentence pair can be encoded first. When encoding, first, according to a preset encoding vocabulary table that provides the mapping relationship data between words and vectors, encode the candidate keyword and the title text respectively to obtain the embedding vector of the candidate keyword and the embedding vector of the title text. Among them, the title text is first segmented to determine all the segments, and then encoded based on the encoding vocabulary table. Finally, the embedding vector of the candidate keyword and the embedding vector of the title text are concatenated into a sentence vector, which can be used for input into the text classification model for classification and discrimination.

[0081] The text classification model includes a text feature extraction network suitable for performing representation learning on sentence vectors corresponding to text information to extract its deep semantic information, and a classifier for performing classification mapping on the deep semantic information output by the text feature extraction network.

[0082] The text feature extraction network can adopt neural network models based on deep learning such as Bert, AlBert, RoBERTa, RoBERTa-wwm-ext, TextCNN, LSTM, etc. The classifier performs a full connection on the deep semantic information through a fully connected layer and maps it to its output layer. In the output layer, a classification function such as Softmax() is used to calculate the classification probabilities of each different relevant category mapped to a preset classification space. Thus, the relevant degree between the candidate keyword and the title text in the sentence pair is characterized by the different relevant categories, thereby determining the information contribution level of the candidate keyword to the title text.

[0083] The number of relevant categories in the classification space can be flexibly configured during the training stage of the text classification model. For example, the classification space contains three relevant categories, representing fully relevant, partially relevant, and irrelevant respectively. In other variants, two or more than three categories can also be set, which can be flexibly determined according to actual needs during the training stage of the text classification model. Among them, when set to two categories, the Sigmoid() function can also be used in the output layer to calculate the classification probability.

[0084] The text classification model is pre-trained until it converges, so that it learns to discriminate the relevant categories suitable for classifying the sentence vectors corresponding to the sentence pair and then is put into use. It is not difficult to understand that according to the principle of the neural network model based on deep learning, the requirements for the input form of the text classification model are the same during the training stage and the inference stage after being put into use.

[0085] Thus, when a sentence pair is encoded as a sentence vector and input into the text classification model, the text classification model extracts deep semantic information from the sentence vector through its text feature extraction network, and then the classifier performs classification mapping according to the deep semantic information, maps it to each relevant category, and obtains the classification probabilities corresponding to each relevant category. The relevant category with the largest classification probability is the relevant category corresponding to the sentence pair determined by the text classification model, which characterizes the relevant degree between the candidate keyword in the sentence pair and the title text, reaches the level indicated by the relevant category, and also characterizes that the information contribution value of the candidate keyword to the title text reaches the information contribution level corresponding to the relevant category.

[0086] After making classification judgments on each sentence pair corresponding to each candidate keyword, the mapping relationship data between each candidate keyword and its related category determined by the text classification model can be obtained.

[0087] Step S1400: Screen out the sentence pairs whose related category is the target related category, and use the candidate keyword therein as the target keyword of the title text.

[0088] As mentioned above, the number of related categories configured in the classification space may include multiple, indicating different degrees of relevance. When it is necessary to screen the candidate keywords, the target related category can be set according to the required degree of relevance. For example, two related categories, "fully relevant" and "partially relevant", can be set as the target related categories for screening the candidate keywords. Then, the candidate keywords whose related category belongs to the target related category are determined as the target keywords, screened out to form a relevant set, and the extraction of the target keywords of the title text is completed.

[0089] In a variant embodiment, in order to facilitate the distinction between product words and attribute words, two sets of candidate keywords belonging to product words and attribute words can be constructed respectively, and a target product word list and a target attribute word list can be obtained respectively.

[0090] The other word segments and other candidate keywords in the title text are substantially filtered out because they are not determined as target keywords and are no longer used. Therefore, the filtering of redundant information in the title text is achieved.

[0091] According to the above embodiments, compared with the prior art, the advantages of the present application include: according to the characteristic that the commodity title is composed of multiple words stacked, sentence pairs are formed by using the title text of the commodity title and the candidate keywords in the title text, and then a text classification model with the ability to classify semantic relevance learned in advance is used to classify and judge each sentence pair to determine the related category between the candidate keyword and the title text. Finally, the candidate keywords in the title text are selected as the target keywords according to the preset target related category. Thus, by using the semantic recognition ability obtained by training the model, the filtering of redundant information in the title text is achieved, and the target keywords with high information value are hit, which is applicable to providing the basic materials required for matching commodities in e-commerce platforms for commodity search, commodity advertising, and commodity recommendation, thereby improving the commodity information matching efficiency of the entire e-commerce platform.

[0092] Please refer to Figure 2 , in the in-depth partial embodiments, step S1200: Extract candidate keywords from the title text, including the following steps:

[0093] Step S1210: Match the title text with a preset product vocabulary to obtain the product words in the title text:

[0094] One exemplary method is to use a rule-based matching approach. Based on tokenizing the title text into multiple tokens, each token is precisely matched with the preset product vocabulary, and the tokens that match the product vocabulary are determined as product words.

[0095] Another exemplary method is also based on a rule-based matching approach. Each token in the product vocabulary is used to search for a corresponding string in the title text in an exact matching manner. When a token in the product vocabulary matches the title text, that token constitutes the corresponding product word in the title text.

[0096] A third exemplary method is based on a semantic matching rule. Based on tokenizing the title text, the semantic vectors of each token are calculated for similarity with the semantic vectors of each token in the product vocabulary. The token in the product vocabulary with the highest similarity exceeding a preset threshold is selected as the product word corresponding to the respective token. When calculating the similarity, any one of the data distance algorithms can be used, including but not limited to the cosine similarity algorithm, Euclidean distance algorithm, Pearson correlation coefficient algorithm, Jaccard coefficient algorithm, etc.

[0097] The described product vocabulary can be pre-extracted by those skilled in the art from a large number of pre-collected product titles to form the prior knowledge for determining the product words in the title text of this application.

[0098] Step S1220: Match the title text with a preset attribute vocabulary to obtain the attribute words in the title text:

[0099] One exemplary method is to use a rule-based matching approach. Based on tokenizing the title text into multiple tokens, each token is precisely matched with the preset attribute vocabulary, and the tokens that match the attribute vocabulary are determined as attribute words.

[0100] Another exemplary method is also based on a rule-based matching approach. Each token in the attribute vocabulary is precisely matched with the title text. When a token in the attribute vocabulary matches the title text, that token constitutes the corresponding attribute word in the title text.

[0101] In the third exemplary method, based on semantic matching rules, after segmenting the title text, the semantic vectors of each segment are used to calculate the similarity with the semantic vectors of each entry in the attribute word library. The entry in the attribute word library with the highest similarity exceeding a preset threshold is selected as the attribute word corresponding to the corresponding segment. When calculating the similarity, any data distance algorithm can be used, including but not limited to the cosine similarity algorithm, Euclidean distance algorithm, Pearson correlation coefficient algorithm, Jaccard coefficient algorithm, etc.

[0102] The described attribute word library can be pre-extracted by those skilled in the art from a large number of pre-collected product titles to form prior knowledge for determining the attribute words of the title text in this application.

[0103] Step S1230: Determine the product word and the attribute word as candidate keywords for the title text.

[0104] After determining the product word and the attribute word from the title text, a product word list and an attribute word list can be created correspondingly. The product word matched from the title text is stored in the product word list as a candidate keyword therein, and the attribute word matched from the title text is stored in the attribute word list as a candidate keyword therein.

[0105] According to the above embodiments, the product words and attribute words of the title text can be determined based on the pre-constructed prior knowledge, that is, determined according to the corresponding product word library and attribute word library. Since the product word library and the attribute word library are effective data processed on the basis of big data and have the function of indicating the knowledge value of each corresponding entry in the title text, thus, the product words and attribute words can be initially extracted from the title text to filter other types of entries in the title text, removing redundant information in the title text and completing data cleaning for subsequent processing of the information value of the entries.

[0106] In an extended partial embodiment, before the step S1230 of determining the product word and the attribute word as candidate keywords for the title text, the following steps are included:

[0107] Step S2100: Filter the product words or attribute words with a semantic similarity lower than a preset threshold according to the semantic similarity between the title text and its product words or attribute words.

[0108] After matching the product words and attribute words from the title text according to the above embodiment, the product words and attribute words can be further selected to determine the candidate keywords. An exemplary selection principle can be determined based on the semantic matching of the product words, attribute words and the title text. In one embodiment, the product words, attribute words and the title text can be encoded one-to-one into an embedding vector with reference to a preset encoding word table. The embedding vector can be used in this embodiment to calculate semantic similarity, and can also be used in the subsequent steps of this application to construct the sentence vector of the input sentence pair of the text classification model to simplify business logic and improve execution efficiency.

[0109] After obtaining the embedding vectors of the title text and each product word and attribute word therein, the semantic similarity between the embedding vectors of each product word, attribute word and the title text is calculated. When calculating the semantic similarity, any data distance algorithm can be used, including but not limited to any one of the cosine similarity algorithm, the Euclidean distance algorithm, the Pearson correlation coefficient algorithm, the Jaccard coefficient algorithm, etc. After determining the semantic similarity corresponding to each product word and attribute word, distinguish between product words and attribute words, and sort the product words and attribute words according to the semantic similarity. Then, based on the product words and attribute words whose semantic similarity exceeds the preset threshold, they are determined as candidate keywords of the title text, and the corresponding product word list and attribute word list are constructed respectively.

[0110] As a variation of some of the embodiments described above, in the process of matching product words and attribute words from the title text, the semantic similarity between the product words, attribute words and the title text has been calculated, and the product words and attribute words have been selected according to a preset threshold. In this case, the process of this step is actually integrated into the process of matching product words and attribute words from the title text, and therefore, there is no need to repeat it.

[0111] According to the above embodiments, after determining product words and attribute words from the title text, the semantic similarity between product words, attribute words and title text can be used to preferentially determine candidate keywords to ensure that the final determined candidate keywords have a higher information contribution value.

[0112] In some extended embodiments, before the step S1300, which forms sentence pairs with each candidate keyword and the title text, and classifies each sentence pair using a text classification model that has been trained to a convergent state, the following steps are included:

[0113] Step S3100: Iteratively train the text classification model using training samples in a preset data set until it reaches a convergence state, wherein the training samples include the title text and a single candidate keyword contained in the title text:

[0114] A preset data set can be used to train the text classification model. By calling the training samples in the data set, the text classification model is iteratively trained until it converges.

[0115] The training samples in the data set can be constructed by collecting a large number of product titles from the product database of an e-commerce platform. A sentence pair is formed by combining the title text of each product title and one of its product words or attribute words as a training sample, that is, a title text and a candidate keyword form a sentence pair, and the candidate keyword is a product word or attribute word in the title text. Correspondingly, the degree of relevance is labeled for the training sample, that is, the corresponding relevant category is set, so as to be used as the supervision label after the text classification model classifies and maps the sentence pair of the training sample to predict its relevant category.

[0116] During training, the sentence pair of each training sample is encoded into a sentence vector according to the aforementioned encoding principle and input into the text classification model. The text feature extraction network in it performs representation learning to obtain corresponding deep semantic information. Then, the classifier performs classification mapping to obtain the classification probabilities corresponding to each category mapped to the classification space, so as to determine the relevant category corresponding to the sentence pair as the relevant category with the highest classification probability.

[0117] After the text classification model predicts the relevant category of a training sample, the cross-entropy function is applied, and the loss value of the relevant category predicted by the model is calculated using the supervision label corresponding to the training sample. Then, it is judged whether the loss value reaches a preset threshold. If it reaches the preset threshold, it indicates that the model has converged and the training can be terminated. Otherwise, the weights of the model are corrected by backpropagation according to the loss value to achieve gradient update, prompting the model to further converge, and then the next training sample is called to continue iterative training. By iteratively training like this, the model can finally reach the convergence state.

[0118] In the above embodiments, by combining the title text and one of its candidate keywords into a sentence pair to form the corresponding training sample for training the text classification model, and using the corresponding supervision label to correct the model loss, the model can be made to converge through iterative training and be used to determine the relevant category of the candidate keyword in the sentence pair.

[0119] It can be seen from this that since the training samples required for training the text classification model can be collected from the product titles in the product database of the e-commerce platform itself, and the title text and the candidate keyword in it are associated as a sentence pair to form the input of the model, it is convenient for the text feature extraction network of the model to extract the association information between the two during the representation learning process, making the model easier to train, thereby reducing the training cost and enabling the model to converge more quickly.

[0120] After each title text obtains its corresponding target keywords such as product words and attribute words, these target keywords themselves carry certain features. For example, the position features corresponding to the positions where the target keywords appear in the title text, and the word frequency features obtained by the target keywords in the preset knowledge system, etc. The position features and word frequency features corresponding to each target keyword constitute its own statistical features. According to each specific feature in the statistical features of each target keyword, it is quantified into a numerical value according to a preset formula, and the numerical values of each specific feature are summarized to determine the information score corresponding to each target keyword. Among them, the summarization method can be direct summation or weighted summation, which can be flexibly implemented by those skilled in the art.

[0121] It can be seen that the information score corresponding to each target keyword is a quantitative representation of the information value of the target keyword. After each target keyword has determined its corresponding information score, the summarized information score corresponding to any combination text between the product word and the attribute word of the title text can be determined according to the information score.

[0122] According to the above principle, please refer to Figure 3 , in some extended embodiments, after the step S1400 of using the candidate keywords as the target keywords of the title text, the following steps are included:

[0123] Step S1500: Determine the word frequency feature of each target keyword of the title text according to the statistical word frequency of each target keyword of the title text in the preset title library:

[0124] The word frequency feature of each target keyword determined from the title text can be determined by referring to the prior knowledge provided by the preset title library. Specifically, at least one title library is prepared, and a large number of product titles of commodities are collected in the title library. Further, on the basis of word segmentation of each commodity title in the title library, the statistical word frequency corresponding to each word segment is counted, and then the mapping relationship data between the word segment and its word frequency is constructed into a word frequency statistical table.

[0125] Thus, when determining the word frequency feature corresponding to each target keyword in the title text, the statistical word frequency of the word segment corresponding to the target keyword is called from the word frequency statistical table, and it is normalized into a corresponding numerical value according to a preset normalization method, and the construction of the word frequency features of each target keyword of the title text can be completed.

[0126] The word frequency features of each target keyword can be manifested as one or more. Specifically, different word frequency statistical tables can be determined through different types of title libraries, and the statistical word frequencies of each target keyword in different title libraries can be obtained respectively, so as to determine multiple word frequency features for each target keyword. When the title library consists of different types of product titles, the provided statistical word frequencies will represent different reference information values. Thus, the content of the word frequency features of each target keyword is enriched from multiple dimensions.

[0127] Step S1600: Determine the position feature of each target keyword in the title text according to its position in the title text:

[0128] The positions where each target keyword appears in the title text are different, including both its absolute position and its relative position to a target keyword belonging to product words, and they are all different. Thus, by applying a preset normalization method to quantify the position information of each target keyword in the title text into a numerical value, the position information can include absolute position information, or relative position information, or both the absolute position information and the relative position information at the same time, and the corresponding position features of each target keyword can be obtained.

[0129] Step S1700: Quantify the word frequency feature and the position feature of each target keyword in the title text to determine the information score of the target keyword:

[0130] As mentioned above, the word frequency features and the position features of each target keyword in the title text have been determined and can be normalized into numerical features. Therefore, for each target keyword, by adding or weighted summing its various word frequency features and various position features, the corresponding information score of the target keyword can be obtained. Thus, each target keyword in the title text, whether it belongs to product words or attribute words, can obtain its corresponding information score. It is not difficult to understand that this information score comprehensively reflects the value of the position information and the word frequency information of the target keyword, and can effectively measure the contribution degree of the target keyword to the information value of the title text.

[0131] According to the above process, by referring to prior knowledge to determine the word frequency features of each target keyword in the title text, and by referring to the position information of each target keyword in the title text to determine its corresponding position feature, and quantifying the information value corresponding to each target keyword according to the word frequency feature and the position feature of each target keyword, the effective quantification of the information value of the target keyword in the title text is realized, providing key reference information for generating the abstract of the title text subsequently.

[0132] Step S1800: Select the combined text of the product word and the attribute word as the title abstract of the commodity according to the information score:

[0133] According to the expression habits of human language, usually one product word can indicate a commodity. By attaching one or more attribute words to the product word, the commodity can be distinguished from other commodities with the same name. Therefore, according to this principle, the product words and attribute words determined as target keywords in the title text can be combined, and one or more combined texts can be selected based on the overall situation of the information scores of the product words and attribute words in the combined text, which can be used as the title abstract of the commodity.

[0134] According to the above embodiments, compared with the prior art, in this application, according to the characteristic that the commodity title is composed of multiple piled-up words, the statistical features are first determined respectively for the product words and attribute words determined as target keywords in the title text of the commodity. Based on the statistical features, the information scores of the product words and attribute words are calculated. Then, according to the information scores, some combinations are selected from the combinations of the product words and attribute words as the abstract of the title text of the commodity. Without applying high-cost machine learning models and deep learning models, the title abstract can be obtained based on the target keywords, with low cost and high operation efficiency. It is especially suitable for providing the basic materials required for matching commodities in commodity search, commodity advertising, and commodity recommendation in the e-commerce platform, thereby improving the commodity information matching efficiency of the entire e-commerce platform.

[0135] Please refer to Figure 4 In some deepened embodiments, step S1500: Determine the word frequency feature of each target keyword in the title text according to the statistical word frequency in the preset title library, including the following steps:

[0136] Step S1510: Determine the first word frequency feature of each target keyword according to the statistical word frequency in the first title library, where the first title library is a title library composed of the title texts of commodities of the same category as the commodity:

[0137] To determine the information value of each target keyword in the first aspect according to the first prior knowledge, the corresponding first word frequency statistical table can be obtained based on the first title library. In the first title library, the commodity titles of the commodities corresponding to the title text in the same category in the commodity classification system are collected in advance, and then, on the basis of word segmentation of the first title library, the same word frequency statistics as described above are performed to obtain the corresponding first word frequency statistical table. Further, referring to the previous embodiment, according to the statistical word frequencies in the first word frequency statistical table, the first word frequency features corresponding to the respective target keywords of the title text can be determined by normalization.

[0138] One of the exemplary normalization methods is to perform logarithmic transformation and min-max normalization on the statistical word frequency, so as to obtain the corresponding first word frequency feature.

[0139] In a modified embodiment, based on the first title library, a corresponding first word frequency statistical table can be constructed for product words and attribute words respectively, so as to quickly obtain the word frequencies of the respective target keywords by referring to them.

[0140] Step S1520: Determine the second word frequency feature according to the statistical word frequency of each target keyword in the second title library, where the second title library is a title library composed of the title texts of products belonging to the same online store as the commodity:

[0141] To determine the information value of each target keyword in the second aspect according to the second prior knowledge, a corresponding second word frequency statistical table can be obtained based on the second title library. In the second title library, the product titles of products belonging to the same online store as the title text are pre-collected, and then the word frequency statistics are performed in the same way as described above on the basis of word segmentation of the second title library, so as to obtain the corresponding second word frequency statistical table. Further, referring to the previous embodiment, according to the statistical word frequency in the second word frequency statistical table, the second word frequency feature corresponding to each target keyword of the title text can be determined by normalization.

[0142] One of the exemplary normalization methods is to perform logarithmic transformation and min-max normalization on the statistical word frequency, so as to obtain the corresponding second word frequency feature.

[0143] In a modified embodiment, based on the second title library, a corresponding second word frequency statistical table can be constructed for product words and attribute words respectively, so as to quickly obtain the word frequencies of the respective target keywords by referring to them.

[0144] According to the above embodiments, by providing corresponding word frequency features for each target keyword in the title text from two dimensions of similar products and products in the same store, usually multiple word frequency features can reflect the information value of the target keyword in similar products and products in the same store, enriching the information reference dimension for determining the information score of the target keyword, overcoming the drawback that the inherent stacking of terms in the product title leads to limited reference information of the terms, and making the information score obtained therefrom more capable of reflecting the comprehensive information value of the corresponding target keyword.

[0145] Please refer to Figure 5 , in some in-depth embodiments, step S1600: Determine the position feature according to the position of each target keyword of the title text in the title text, including the following steps:

[0146] Step S1610: Determine the absolute position feature of each target keyword according to its absolute position in the title text:

[0147] The order in which each target keyword appears in the title text, that is, its absolute position information, serves to represent the position feature of the target keyword. Thus, the arrangement order of each target keyword in the title text can be correspondingly determined, and this arrangement order is normalized into a numerical value using a preset method and determined as its corresponding absolute position feature.

[0148] In an exemplary normalization method, the following formula is applied for processing:

[0149] F abs = 1 / log2(1 + L abs )

[0150] where L abs represents the numerical value corresponding to the arrangement order of the target keyword, and F abs is the absolute position feature obtained after normalization using the formula.

[0151] For each target keyword belonging to a product word and an attribute word, normalization can be performed according to the above principle, so as to determine the absolute position feature corresponding to each target keyword.

[0152] Step S1620: Determine the relative position feature of each target keyword belonging to an attribute word according to its relative position in the title text with respect to its closest product word:

[0153] For the relative position feature of each target keyword in the title text, according to the difference in its part of speech, the processing in this embodiment is different.

[0154] For each attribute word, its relative position feature can be determined by referring to its relative position information with respect to its closest product word. First, determine the difference in its sorting order, and then normalize according to this difference. For example, in the title text "Fashionable and elegant shawl cape knitted woolen sweater", the attribute word "Fashionable" (position serial number 1), which is determined as a target keyword, and the product word "shawl" (position serial number 5), which is the closest and determined as a target keyword after it, the relative position information obtained through the operation of the position serial number difference is 4. After determining this relative position information, the following formula can be applied to normalize it to obtain the relative position feature corresponding to the attribute word:

[0155] F rel = 1 / log2(1 + L rel )

[0156] where L rel represents the numerical value corresponding to the relative position information of the attribute word, and Frel The relative position feature obtained after applying formula normalization.

[0157] Step S1630: Determine the relative position feature of each target keyword belonging to the product word with a standard value:

[0158] For the target keyword belonging to the product word, since its closest product word is itself, if the method of referring to the attribute word is used to determine the corresponding value of its relative position information, a value of 0 will be obtained and there will be no feature representation value. Therefore, a standard value, such as the numerical value 1, can be used to uniformly describe the relative position information of the product word. Then, in the same way as the attribute word, the normalization formula is applied to normalize this standard value to obtain the corresponding relative position feature of the product word.

[0159] The absolute position feature and the relative position feature of each target keyword together constitute the position feature of the target keyword. It is not difficult to understand that the position information of the target keyword in the title text and the relative distance information between each target keyword and the product word in the title text are jointly described through the absolute position feature and the relative position feature, enriching the statistical features of each target keyword and further realizing the effective representation of the information value of each target keyword.

[0160] Please refer to Figure 6 , in some deepened embodiments, the step S1800: Select the combined text of the product word and the attribute word as the title abstract of the commodity according to the information score, including the following steps:

[0161] Step S1811: Screen out the product words and attribute words determined as target keywords with information scores higher than the preset threshold:

[0162] After each product word and attribute word determined as a target keyword in the title text obtains its corresponding information score, the corresponding information value of each target keyword is reflected. Therefore, according to the information scores corresponding to each target keyword, a preset threshold is provided in advance, which can be used to realize the optimization of each product word and attribute word. Since the information scores of the product word and the attribute word are both normalization results, the preset thresholds for the product word and the attribute word can be the same preset threshold, and this preset threshold can be a measured threshold or an empirical threshold, which can be flexibly set by those skilled in the art.

[0163] Accordingly, select the product words and attribute words with information scores higher than the preset threshold respectively to form new product word tables and attribute word tables for calling.

[0164] Step S1812: Combine any number of the screened attribute words with any single screened product word to construct the corresponding title abstract:

[0165] To obtain a title abstract, any product term can be selected from the described product vocabulary, and any number of one or more attribute terms can be selected from the attribute vocabulary. According to the natural language expression habit, the product term is placed after, and the attribute terms are placed before, and they are combined arbitrarily to construct multiple combined texts, and these combined texts constitute the title abstract of the title text of the commodity.

[0166] According to the embodiments herein, through information scoring, the product terms and attribute terms determined as target keywords in the title text of the commodity are preferentially selected, and the target keywords with higher information value are selected. Then, these target keywords are flexibly combined according to the natural language expression habit to obtain the corresponding title abstract. The obtained title abstract is a string composed of target keywords with higher information value, and the redundant information in the title text has been filtered. Moreover, the attribute terms and product terms have been filtered in advance, which plays a role of outlining the title text and has the function of concisely indicating the title text.

[0167] Please refer to Figure 7 , in some deepened embodiments, the step S1400 of selecting the combined text of the product term and the attribute term as the title abstract of the commodity according to the information scoring includes the following steps:

[0168] Step S1821: Combine any number of attribute terms determined as target keywords with a single product term determined as a target keyword to construct a corresponding candidate title abstract:

[0169] In this embodiment replaced with the previous embodiment, according to the natural language expression habit, the product terms determined as target keywords obtained in the title text can be placed after, and the attribute terms determined as target keywords among them can be placed before, and they are combined arbitrarily to construct multiple combined texts in advance, and these combined texts constitute the candidate title abstract of the title text of the commodity. Similarly, in the combined text, generally only a single product term is included, but one or more attribute terms can be included.

[0170] Step S1822: Screen out the candidate title abstracts with a total score value higher than the preset threshold from the total score values of the information scores of each target keyword in the candidate title abstracts as the finally determined title abstract:

[0171] In the candidate title abstracts, the corresponding information scores of each attribute term and product term have been determined in advance. At this point, the information scores of each attribute term and product term can be summarized, that is, addition or weighted addition operations are performed to obtain the total score value corresponding to each candidate title abstract.

[0172] On the basis that each candidate title abstract has obtained its corresponding total score, the candidate title abstracts are sorted in reverse order according to the total score. Then, one or more candidate title abstracts with the top rankings are intercepted according to a preset quantity, and used as the finally determined title abstract corresponding to the title text.

[0173] Alternatively, in an alternative manner, a preset threshold is compared with the total scores of the candidate title abstracts, and the candidate title abstracts with total scores higher than the preset threshold can be determined as the final title abstracts. Since the information scores of the product words and attribute words are both normalized results, the preset thresholds for the product words and attribute words can be the same preset threshold, and this preset threshold can be a measured threshold or an empirical threshold, which can be flexibly set by those skilled in the art.

[0174] According to the above embodiments, by first combining the attribute words and product words of the title text according to the natural language expression habits to obtain multiple combined texts, obtaining candidate title abstracts, and then sorting and selecting the candidate title abstracts according to the total scores obtained by summing up the information scores of each target keyword in each candidate title abstract, an effective title abstract can be determined from the idea of the best combination of product words and attribute words. Since the existence of multiple attribute words often leads to a higher total score of the entire title abstract than when there are fewer attribute words, relatively speaking, the title abstract determined in this way is closer to the original title text and provides richer information.

[0175] Please refer to Figure 8 , a device for extracting keywords of a commodity title provided to meet one of the purposes of this application, is a functional embodiment of the method for extracting keywords of a commodity title of this application. The device includes a title acquisition module 1100, a term extraction module 1200, a term classification module 1300, and a target determination module 1400, where: the title acquisition module 1100 is used to acquire the title text of the commodity; the term extraction module 1200 is used to extract candidate keywords belonging to product words and attribute words from the title text; the term classification module 1300 is used to form sentence pairs with each candidate keyword and the title text, and use a text classification model that has been trained to a convergent state to classify each sentence pair respectively, and determine the relevant category representing the relevance between the candidate keyword in each sentence pair and the title text; the target determination module 1400 is used to screen out the sentence pairs with the relevant category being the target relevant category, and use the candidate keywords therein as the target keywords of the title text.

[0176] In a further partial embodiment, the entry extraction module 1200 includes: a product word extraction unit configured to match the title text with a preset product word library to obtain product words in the title text; an attribute word extraction unit configured to match the title text with a preset attribute word library to obtain attribute words in the title text; and a candidate set determination unit configured to determine the product words and the attribute words as candidate keywords for the title text.

[0177] In an extended partial embodiment, prior to the candidate set determination unit, it includes: a similarity filtering unit configured to filter product words or attribute words with a semantic similarity lower than a preset threshold according to the semantic similarity between the title text and its product words or attribute words.

[0178] In an extended partial embodiment, prior to the entry classification module 1300, it includes: a model training module configured to perform iterative training on a text classification model using training samples in a preset dataset until it converges, where the training samples include title texts and individual candidate keywords included in the title texts.

[0179] In an extended partial embodiment, after the target determination unit, it includes: a word frequency feature determination unit configured to determine the word frequency feature of each target keyword in the title text according to the statistical word frequency of the target keyword in a preset title library; a position feature determination unit configured to determine the position feature of each target keyword in the title text according to the position of the target keyword in the title text; an information score determination unit configured to quantify the word frequency feature and the position feature of each target keyword in the title text to determine the information score of the target keyword; and a title abstract selection unit configured to select the combined text of the product words and the attribute words as the title abstract of the product according to the information score.

[0180] In a further partial embodiment, the word frequency feature determination unit includes: a first word frequency feature unit configured to determine the first word frequency feature of each target keyword according to the statistical word frequency of the target keyword in a first title library, where the first title library is a title library composed of title texts of products of the same category as the product; and a second word frequency feature unit configured to determine the second word frequency feature of each target keyword according to the statistical word frequency of the target keyword in a second title library, where the second title library is a title library composed of title texts of products in the same online store as the product.

[0181] In some of the in - depth embodiments, the position feature determination unit includes: an absolute position feature unit for determining the absolute position feature of each target keyword according to its absolute position in the title text; an attribute word relative position unit for determining the relative position feature of each target keyword belonging to the attribute word according to its relative position relative to the closest product word in the title text; and a product word relative position unit for determining the relative position feature of each target keyword belonging to the product word with a standard value.

[0182] To solve the above - mentioned technical problems, an embodiment of the present application also provides a computer device. As Figure 9 shown, it is a schematic internal structure diagram of the computer device. The computer device includes a processor, a computer - readable storage medium, a memory, and a network interface connected through a system bus. Among them, the computer - readable storage medium of the computer device stores an operating system, a database, and computer - readable instructions. The database can store a control information sequence. When the computer - readable instructions are executed by the processor, the processor can implement a method for extracting keywords from a product title. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device can store computer - readable instructions. When the computer - readable instructions are executed by the processor, the processor can execute the method for extracting keywords from a product title of the present application. The network interface of the computer device is used to connect and communicate with a terminal. Those skilled in the art can understand that Figure 9 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0183] In this embodiment, the processor is used to execute Figure 8 the specific functions of each module and its sub - modules in . The memory stores the program code and various types of data required to execute the above - mentioned modules or sub - modules. The network interface is used for data transmission between the user terminal and the server. The memory in this embodiment stores the program code and data required to execute all modules / sub - modules in the device for extracting keywords from a product title of the present application. The server can call the program code and data of the server to execute the functions of all sub - modules.

[0184] The present application also provides a storage medium storing computer - readable instructions. When the computer - readable instructions are executed by one or more processors, one or more processors are caused to execute the steps of the method for extracting keywords from a product title according to any embodiment of the present application.

[0185] The present application also provides a computer program product, including a computer program / instructions, which, when executed by one or more processors, implement the steps of the method according to any embodiment of the present application.

[0186] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments of the present application can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a computer-readable storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0187] In summary, the present application utilizes the semantic recognition ability obtained through training of the text classification model to filter redundant information in the title text, and hits target keywords with high information value, which is applicable to providing basic materials required for matching products in product search, product advertising, and product recommendation in an e-commerce platform, thereby improving the product information matching efficiency of the entire e-commerce platform.

[0188] Those skilled in the art of the present technology can understand that the steps, measures, and solutions in the various operations, methods, and processes discussed in the present application can be alternated, changed, combined, or deleted. Further, the other steps, measures, and solutions in the various operations, methods, and processes discussed in the present application can also be alternated, changed, rearranged, decomposed, combined, or deleted. Further, the steps, measures, and solutions in the prior art that are the same as those disclosed in the various operations, methods, and processes in the present application can also be alternated, changed, rearranged, decomposed, combined, or deleted.

[0189] The above are only partial embodiments of the present application. It should be noted that for those of ordinary skill in the art of the present technology, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A method for extracting keywords from a product title, characterized in that, It includes the following steps: Obtain the title text of the product; Extract candidate keywords belonging to product words and attribute words from the title text; Form sentence pairs by combining each candidate keyword with the title text, and use a text classification model that has been trained to a convergent state to classify each sentence pair respectively, and determine the relevant category representing the relevance degree of the candidate keyword in each sentence pair to the title text; Filter out the sentence pairs whose relevant category is the target relevant category, and use the candidate keywords therein as the target keywords of the title text; Determine the word frequency feature of each target keyword in the title text according to the statistical word frequency of each target keyword of the title text in a preset title library, including: determining its first word frequency feature according to the statistical word frequency of each target keyword in the first title library, and the first title library is composed of the title texts of products of the same category as the product; determining its second word frequency feature according to the statistical word frequency of each target keyword in the second title library, and the second title library is composed of the title texts of products in the same online store as the product; Determine the position feature of each target keyword in the title text according to its position in the title text, including: determining its absolute position feature according to the absolute position of each target keyword in the title text; determining its relative position feature according to the relative position of each target keyword belonging to the attribute word in the title text relative to its closest product word; determining its relative position feature with a standard value for each target keyword belonging to the product word; Quantitatively determine the information score of each target keyword according to the word frequency feature and position feature of each target keyword in the title text; Select the combined text of the product word and the attribute word as the title abstract of the product according to the information score; 2. The method for extracting keywords of a product title according to claim 1, wherein Extract candidate keywords from the title text, including the following steps: Match the title text with a preset product word library to obtain the product words in the title text; Match the title text with a preset attribute word library to obtain the attribute words in the title text; Determine the product word and the attribute word as the candidate keywords of the title text; 3. The method for extracting keywords of a product title according to claim 2, wherein Before the step of determining the product word and the attribute word as the candidate keywords of the title text, it includes the following steps: Filter out the product words or attribute words whose semantic similarity is lower than a preset threshold according to the semantic similarity between the title text and its product words or attribute words; 4. The method for extracting keywords of a product title according to claim 3, wherein Before the step of forming sentence pairs by combining each candidate keyword with the title text and using a text classification model that has been trained to a convergent state to classify each sentence pair respectively, it includes the following steps: Perform iterative training on the text classification model using the training samples in a preset dataset, and train it to a convergent state, and the training samples include the title text and a single candidate keyword included in the title text; 5. A device for extracting keywords of a product title, characterized in that, It includes: A title acquisition module for acquiring the title text of the product; A term extraction module for extracting candidate keywords belonging to product words and attribute words from the title text; The term classification module is used to form sentence pairs by combining each candidate keyword with the title text, and respectively classify each sentence pair using a text classification model that has been trained to a convergent state, so as to determine the relevant category representing the relevance between the candidate keyword in each sentence pair and the title text; The target determination module is used to filter out the sentence pairs whose relevant category is the target relevant category, and use the candidate keywords therein as the target keywords of the title text; The word frequency feature determination unit is used to determine the word frequency feature of each target keyword of the title text according to the statistical word frequency of each target keyword of the title text in a preset title library, including: determining the first word frequency feature according to the statistical word frequency of each target keyword in the first title library, where the first title library is a title library composed of the title texts of commodities of the same category as the commodity; determining the second word frequency feature according to the statistical word frequency of each target keyword in the second title library, where the second title library is a title library composed of the title texts of commodities in the same online store as the commodity; The position feature determination unit is used to determine the position feature of each target keyword of the title text according to the position of each target keyword of the title text in the title text, including: determining the absolute position feature according to the absolute position of each target keyword in the title text; determining the relative position feature according to the relative position of each target keyword belonging to the attribute word in the title text relative to its closest product word; determining the relative position feature of each target keyword belonging to the product word with a standard value; The information score determination unit is used to quantitatively determine the information score of each target keyword of the title text according to the word frequency feature and position feature of each target keyword of the title text; The title abstract selection unit is used to select the combined text of the product word and the attribute word as the title abstract of the commodity according to the information score; 6. The apparatus for extracting keywords of a product title according to claim 5, wherein, The term extraction module includes: The product word extraction unit is used to match the title text with a preset product word library to obtain the product words in the title text; The attribute word extraction unit is used to match the title text with a preset attribute word library to obtain the attribute words in the title text; The candidate set determination unit is used to determine the product word and the attribute word as the candidate keywords of the title text; 7. The device for extracting keywords of a product title according to claim 6, wherein The term extraction module further includes: The similarity filtering unit is used to filter out the product words or attribute words whose semantic similarity is lower than a preset threshold according to the semantic similarity between the title text and its product words or attribute words; 8. The apparatus for extracting commodity title keywords according to claim 7, wherein, This device further includes: The model training module is used to perform iterative training on the text classification model using the training samples in a preset dataset and train it to a convergent state, where the training samples include the title text and the single candidate keyword included in the title text; 9. A computer device, comprising a central processing unit and a memory, characterized in that, The central processing unit is used to call and run the computer program stored in the memory to execute the steps of the method described in any one of claims 1 to 4.

10. A computer-readable storage medium, characterized in that, It stores a computer program implemented by the method according to any one of claims 1 to 4 in the form of computer-readable instructions. When the computer program is called and run by a computer, it executes the steps included in the corresponding method.

Citation Information

Patent Citations

  • Keyword extraction method

    CN103399901A

  • Text key information identification method, electronic apparatus and readable storage medium

    CN108664473A

  • Commodity short title generation method and device

    CN111191022A