Product title abstract generation method and apparatus, device, and medium
By extracting and scoring product terms and attribute terms from product titles, and combining this with a text classification model, high-quality product summaries are generated, solving the problem of redundant information in e-commerce platforms and improving the efficiency of product information matching.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BUSINESS LINE COMMERCIAL PTE LTD
- Filing Date
- 2022-08-12
- Publication Date
- 2026-04-24
AI Technical Summary
Product titles on e-commerce websites contain redundant and noisy information, and existing natural language processing technologies cannot effectively extract key summaries.
By acquiring product title text, extracting product terms and attribute terms, calculating information scores, constructing a candidate summary set, and using a pre-trained text classification model to select high-quality summaries.
Generate high-quality product title summaries to improve the efficiency of product information matching on e-commerce platforms, especially suitable for product search, advertising, and recommendations.
Smart Images

Figure CN115203400B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method for generating product title summaries and the corresponding apparatus, computer equipment, and computer-readable storage medium. Background Technology
[0002] Products sold on e-commerce websites often have titles filled with descriptive terms to boost SEO (Search Engine Optimization) traffic, resulting in overly long titles containing redundant or even irrelevant information. In algorithms for product search, advertising, and recommendation, product titles are crucial input; it's essential to filter out redundancy and noise and extract key summary information.
[0003] Natural language processing technology offers solutions for extracting keywords from a short or long text, but these solutions do not take into account the specific characteristics of product titles on e-commerce websites, and therefore cannot be directly used to process product titles.
[0004] Product titles are typically sentences filled with numerous words and lacking complete grammatical structure, which differs significantly from the text processed by existing keyword extraction technologies. Therefore, extracting summary information from product titles remains a pressing problem to be solved. Summary of the Invention
[0005] The primary objective of this application is to solve at least one of the aforementioned problems by providing a method for generating product title summaries and the corresponding apparatus, computer equipment, and computer-readable storage medium.
[0006] To achieve the various objectives of this application, the following technical solution is adopted:
[0007] A method for generating a product title summary, provided for one of the purposes of this application, includes the following steps:
[0008] Get the product title text;
[0009] Knowledge terms belonging to product terms and attribute terms are extracted from the title text. The information score of each knowledge term is determined by the statistical characteristics of the knowledge terms. Based on the information score, the corresponding combination text of product terms and attribute terms is selected to construct the first candidate summary set.
[0010] Calculate the similarity between multiple long texts composed of partial word units from the title text and the title text, and select the long texts with higher similarity to construct a second candidate summary set;
[0011] The title text is paired with each candidate summary from the first and second candidate summary sets to form a data pair. This pair is then input into a pre-trained and converged text classification model to predict the quality score corresponding to each candidate summary. The candidate summary with the higher quality score is selected as the summary of the title text.
[0012] On the other hand, a product title summary generation device provided to meet one of the purposes of this application includes a title acquisition module, a first set construction module, a second set construction module, and a summary generation module, wherein: the title acquisition module is used to acquire the title text of the product; the first set construction module is used to extract knowledge terms belonging to product words and attribute words from the title text, determine the information score of each knowledge term based on the statistical characteristics of the knowledge terms, and select the corresponding combination text of product words and attribute words to construct a first candidate summary set based on the information score; the second set construction module is used to calculate the similarity between multiple long texts composed of corresponding combinations of some word elements in the title text and the title text, and select the long texts with higher similarity to construct a second candidate summary set; the summary generation module is used to form data pairs with the title text and each candidate summary in the first and second candidate summary sets, input them into a pre-trained and converged text classification model, predict the quality score corresponding to each candidate summary, and select the candidate summary with higher quality score as the summary of the title text.
[0013] In another aspect, a computer device provided for one of the purposes of this application includes a central processing unit and a memory, the central processing unit being configured to invoke and run a computer program stored in the memory to perform the steps of the product title summary generation method described in this application.
[0014] In another aspect, a computer-readable storage medium is provided to suit another purpose of this application, which stores, in the form of computer-readable instructions, a computer program implemented according to the described product title summary generation method, which, when invoked by a computer, performs the steps included in the method.
[0015] The technical solution of this application has many advantages, including but not limited to the following aspects:
[0016] On one hand, based on the characteristic that product titles consist of multiple words, this application first determines the statistical features of product words and attribute words in the product title text. Based on these statistical features, information scores for product words and attribute words are calculated. Then, based on the information scores, a subset of combinations of product words and attribute words is selected as the first candidate summary set for the product title text. On the other hand, the similarity between the title text and multiple long texts composed of corresponding combinations of words from the title text is calculated, and the long texts with higher similarity are selected to construct the second candidate summary set. Further, a pre-trained and converged text classification model is used to predict the quality score corresponding to each candidate summary in the first and second candidate summary sets, and the candidate summaries with higher quality scores are selected as the summary of the title text. It can be seen that by employing two implementation methods to deeply mine the title text, a sufficient number of candidate summaries of considerable quality are initially generated. Based on this, a text classification model is used to accurately predict the quality score corresponding to each candidate summary, thereby selecting high-quality candidate summaries as the summary of the title text.
[0017] On the other hand, the title text summary generated by this application is particularly suitable for providing basic materials needed for matching products in e-commerce platforms for product search, product advertising, and product recommendation, thereby improving the efficiency of product information matching for the entire e-commerce platform. Attached Figure Description
[0018] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0019] Figure 1 A flowchart illustrating a typical embodiment of the product title summary generation method of this application;
[0020] Figure 2 This is a flowchart illustrating the process of obtaining knowledge entries corresponding to product terms and attribute terms in the title text and determining the information score corresponding to each knowledge entry in an embodiment of this application.
[0021] Figure 3 This is a schematic diagram illustrating the process of constructing a second candidate summary set in an embodiment of this application;
[0022] Figure 4 This is a schematic diagram illustrating the training process of the text classification model in an embodiment of this application;
[0023] Figure 5 This is a schematic diagram of the process for preparing the training set in an embodiment of this application;
[0024] Figure 6 A schematic block diagram of the device for generating a product title summary for this application;
[0025] Figure 7This is a schematic diagram of the structure of a computer device used in this application. Detailed Implementation
[0026] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.
[0027] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this application means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.
[0028] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.
[0029] Those skilled in the art will understand that the terms "client," "terminal," and "terminal device" as used herein include both devices that receive wireless signals, devices that only possess wireless signal receiver capabilities without transmission capabilities, and devices with receiving and transmitting hardware, devices that have receiving and transmitting hardware capable of bidirectional communication over a bidirectional communication link. Such devices may include: cellular or other communication devices such as personal computers or tablets, having single-line displays, multi-line displays, or cellular or other communication devices without multi-line displays; PCS (Personal Communications Service) that can combine voice, data processing, fax, and / or data communication capabilities; PDA (Personal Digital Assistant) that may include a radio frequency receiver, pager, internet / intranet access, web browser, notepad, calendar, and / or GPS (Global Positioning System) receiver; and conventional laptops and / or handheld computers or other devices that have and / or include radio frequency receivers. As used herein, "client," "terminal," and "terminal device" can be portable, transportable, installed in a means of transportation (air, sea, and / or land), or suitable and / or configured to operate locally and / or in a distributed manner, operating in any other location on Earth and / or in space. "Client," "terminal," and "terminal device" as used herein can also be a communication terminal, an internet access terminal, or a music / video playback terminal, such as a PDA, a MID (Mobile Internet Device), and / or a mobile phone with music / video playback capabilities, or a smart TV, set-top box, etc.
[0030] The hardware referred to by the names "server," "client," and "service node" in this application is essentially an electronic device with the equivalent capabilities of a personal computer. It is a hardware device with the necessary components revealed by the von Neumann architecture, such as a central processing unit (including an arithmetic logic unit and a control unit), memory, input devices, and output devices. The computer program is stored in its memory, and the central processing unit loads the program stored in the secondary storage into the main memory to run it, execute the instructions in the program, and interact with the input and output devices to complete specific functions.
[0031] It should be noted that the concept of "server" used in this application can also be extended to the case of server clusters. Based on the network deployment principles understood by those skilled in the art, the servers should be logically divided. Physically, these servers can be independent of each other but accessible through interfaces, or they can be integrated into a single physical computer or a computer cluster. Those skilled in the art should understand this flexibility and should not use it to constrain the implementation of the network deployment method in this application.
[0032] One or more of the technical features of this application, unless explicitly specified herein, can be deployed on a server and accessed by a client remotely calling the online service interface provided by the server, or can be directly deployed and run on a client to access the service.
[0033] Unless otherwise specified, the neural network models referenced or potentially referenced in this application may be deployed on a remote server and invoked remotely on the client, or deployed on a client with the capability to invoke directly. In some embodiments, when running on the client, the corresponding intelligence may be acquired through transfer learning in order to reduce the requirements on the client's hardware resources and avoid excessive consumption of the client's hardware resources.
[0034] Unless otherwise specified, all data involved in this application may be stored remotely on a server or on a local terminal device, as long as it is suitable for use by the technical solution of this application.
[0035] Those skilled in the art will understand that although the various methods in this application are described based on the same concept and thus present commonality among them, they can be performed independently unless otherwise specified. Similarly, the various embodiments disclosed in this application are all based on the same inventive concept; therefore, concepts expressed in the same way, as well as concepts that are appropriately changed for convenience but are expressed differently, should be understood equivalently.
[0036] Unless otherwise expressly stated, the various embodiments disclosed in this application can be combined in a cross-cutting manner to flexibly construct new embodiments, as long as such combination does not depart from the inventive spirit of this application and can meet the needs of the prior art or solve a certain deficiency in the prior art. Those skilled in the art should be aware of such modifications.
[0037] The product title summary generation method of this application can be programmed into a computer program product and deployed on a client or server for execution. For example, in an exemplary application scenario of this application, it can be deployed on the server of an e-commerce platform, thereby allowing human-computer interaction with the process of the computer program product through a graphical user interface by accessing the interface opened after the computer program product is run.
[0038] Please see Figure 1 The product title summary generation method of this application, in its typical embodiment, includes the following steps:
[0039] Step S1100: Obtain the product title text;
[0040] To adapt to different specific application scenarios, the product title text can be obtained from the corresponding channels for each application scenario, for example:
[0041] In one exemplary application scenario, when merchants on an e-commerce platform publish product promotion activities, they need to fill in text containing relevant product information. At this time, the summary extracted from the product title can be directly used as material for writing promotional copy.
[0042] In one exemplary application scenario, online merchants on an e-commerce platform need to publish product information for their online goods, including title text corresponding to the product titles. This title text can be collected to generate a summary, which can then be recommended to the merchants.
[0043] In another exemplary application scenario, e-commerce platforms need to determine matching products based on query information provided by consumers, thereby enabling functions such as product search, product advertising, and product recommendation. To this end, corresponding summaries can be generated from the product titles of each product in the product database of the online stores of the e-commerce platform. By matching the query information with the summaries, some products can be determined to be recommended to consumers.
[0044] In another exemplary application scenario, it is necessary to push a summary of a product to relevant users, including consumer users or merchant users. This summary includes the title information of the product. Therefore, a corresponding summary can be generated based on the title text obtained from the product information and encapsulated in the summary information before being pushed to the user.
[0045] In this way, depending on the specific application scenario of the e-commerce platform, the title text can be obtained from multiple sources.
[0046] Step S1200: Extract knowledge terms belonging to product terms and attribute terms from the title text, determine the information score of each knowledge term based on the statistical characteristics of the knowledge terms, and select the corresponding combination text of product terms and attribute terms to construct the first candidate summary set based on the information score;
[0047] The title text typically contains multiple terms, referred to as knowledge terms. These knowledge terms mainly consist of product terms with noun attributes and attribute terms with adjective or adverb attributes. Product terms primarily indicate the name of the product or its equivalent, while attribute terms mainly describe information about the product's characteristics, functions, effects, or other specific attributes. Various methods can be used to extract the knowledge terms corresponding to the product terms and attribute terms from the title text.
[0048] For example, for clothing items, the product terms include, but are not limited to: short-sleeved shirts, T-shirts, polo shirts, coats, and dresses, and the attribute terms include, but are not limited to: long-sleeved, crew neck, casual, wool, and collaboration.
[0049] Extracting knowledge terms from title text is quite flexible. For example, it can be based on rule matching or semantic matching to extract corresponding product terms and attribute terms. Furthermore, in some modified embodiments, the title text can be segmented into words, and then different matching methods can be used on each of the resulting word units to determine whether it belongs to a product term or an attribute term.
[0050] It is understood that the knowledge entries corresponding to the product terms and attribute terms themselves carry certain characteristics. For example, the positional features corresponding to the location of the knowledge entry in the title text, and the word frequency features obtained by the knowledge entry in the preset knowledge system. The positional features and word frequency features corresponding to each knowledge entry constitute its own statistical features. Based on the specific features in the statistical features of each knowledge entry, they are quantified into numerical values according to a preset formula. By summing the values of each specific feature, the information score corresponding to each knowledge entry can be determined. The summarization method can be direct weighting or weighted summation, which can be flexibly implemented by those skilled in the art.
[0051] Typically, a single product term can indicate a product. However, by adding one or more attribute terms to the product term, the product can be distinguished from other products with the same name. Based on this, the product term in the title text is placed at the end and the attribute terms are placed at the beginning. The scores of the information of each product term and attribute term in the combined text are used to select multiple combined texts to construct a first candidate summary set.
[0052] Step S1300: Calculate the similarity between the title text and multiple long texts composed of corresponding combinations of partial words in the title text; select the long texts with higher similarity to construct a second candidate summary set.
[0053] A long text composed of combinations of certain words from the title text can serve as a summary of the title text. A word refers to the smallest unit in the title text capable of independently expressing complete semantics; specifically, it is a single word in the title text. Multiple words in adjacent positions within the title text can be combined using an adjacency combination method. The specific number of words combined can be determined by those skilled in the art as needed. Those skilled in the art should understand that, in practical implementation, the title text can be segmented to obtain the individual words contained within it. Furthermore, based on the foregoing disclosure, a long text can be obtained by combining certain words from the title text using an adjacency combination method.
[0054] It is easy to understand that the long text is equivalent to extracting a portion of the title text. If it is highly similar to the title text semantically, it means that the extracted portion of the text is the key point of the title text, that is, the long text can serve as a summary of the title text to a certain extent.
[0055] Furthermore, the obtained long texts are filtered to select those that are highly semantically similar to the title text. Specifically, a pre-trained convergent text similarity model can be used to vectorize the deep semantic features corresponding to the title text and the multiple long texts. Then, a similarity function is used to calculate the similarity between the title text and the vectorized representations of each long text. It can be understood that the similarity represents the degree of semantic similarity between the title text and each long text; the higher the similarity, the more semantically similar the long text and the title text are. Thus, long texts with similarity exceeding a preset threshold can be selected to construct a second candidate summary set. The specific value of the threshold can be set by those skilled in the art as needed. The text similarity model can be a deep semantic learning-based network model in the field of NLP (Natural Language Processing) suitable for extracting text semantic features. Specifically, the open-source framework Sentence Transformers is used, which provides a large number of pre-trained convergent Transformer models, such as BERT, RoBERTa, XLM-RoBERTa, and MPNet.
[0056] Step S1400: The title text and each candidate summary in the first candidate summary set and the second candidate summary set are used to form a data pair, which is then input into a pre-trained and converged text classification model to predict the quality score corresponding to each candidate summary. The candidate summary with the higher quality score is selected as the summary of the title text.
[0057] To determine whether each candidate summary in the first candidate summary set and the second candidate set is suitable as a summary of the title text, a pre-trained text classification model that has reached convergence can be used for identification.
[0058] The text classification model can be trained under supervision using a training set constructed in advance with labeled training samples. This text classification model can be a model that adds a matching classifier to the text feature extraction model. The training samples are data pairs consisting of the title text of the collected products and the corresponding summary text. The labels indicate whether the summary text of the training sample can be used as a summary of the corresponding title text.
[0059] Specifically, the text classification model extracts features from training samples to obtain a vectorized representation of the semantic features corresponding to the ability of the summary text in the training sample to serve as a summary of the title text. A classifier then classifies this vectorized representation to obtain a classification result predicting the training sample as a positive sample. The positive class indicates that the summary text in the training sample can serve as a summary of the title text. Under the supervision of the label corresponding to the training sample, the model continuously approaches convergence, and this training continues until convergence is achieved. It is easy to understand that after convergence through this supervised training, the text classification model can learn the ability to determine whether the summary in an input data pair can serve as a summary of the title text.
[0060] The title text is concatenated with each candidate summary from the first and second candidate summary sets to form a sentence pair, which is then input into a pre-trained and converged text classification model. The semantic features corresponding to the candidate summary in each data pair that can serve as a summary of the title text are extracted, and a corresponding vectorized representation is generated. A classifier is used to map the vectorized representation to a preset classification space to predict the quality score corresponding to each candidate summary that represents its ability to serve as a summary of the title text. Then, from the first and second candidate summary sets, candidate summaries with quality scores exceeding a preset threshold can be selected as the summary of the title text. The specific value of the threshold can be set as needed by those skilled in the art.
[0061] As can be seen from the typical embodiments of this application, the technical solution of this application has many advantages, including but not limited to the following aspects:
[0062] On one hand, based on the characteristic that product titles consist of multiple words, this application first determines the statistical features of product words and attribute words in the product title text. Based on these statistical features, information scores for product words and attribute words are calculated. Then, based on the information scores, a subset of combinations of product words and attribute words is selected as the first candidate summary set for the product title text. On the other hand, the similarity between the title text and multiple long texts composed of corresponding combinations of words from the title text is calculated, and the long texts with higher similarity are selected to construct the second candidate summary set. Further, a pre-trained and converged text classification model is used to predict the quality score corresponding to each candidate summary in the first and second candidate summary sets, and the candidate summaries with higher quality scores are selected as the summary of the title text. It can be seen that by employing two implementation methods to deeply mine the title text, a sufficient number of candidate summaries of considerable quality are initially generated. Based on this, a text classification model is used to accurately predict the quality score corresponding to each candidate summary, thereby selecting high-quality candidate summaries as the summary of the title text.
[0063] On the other hand, the title text summary generated by this application is particularly suitable for providing basic materials needed for matching products in e-commerce platforms for product search, product advertising, and product recommendation, thereby improving the efficiency of product information matching for the entire e-commerce platform.
[0064] Please see Figure 2 In a further embodiment, step S1200, which involves extracting knowledge terms belonging to product terms and attribute terms from the title text and determining the information score of each knowledge term based on its statistical characteristics, includes the following steps:
[0065] Step S1210: Match the title text with a preset product thesaurus to obtain the knowledge entries belonging to product terms in the title text;
[0066] One exemplary approach is to use rule-based matching to obtain multiple knowledge terms by segmenting the title text, and then accurately match each knowledge term with a preset product terminology database, identifying the knowledge terms that match the product terminology database as product terms.
[0067] Another exemplary approach is to use rule-based matching to find the corresponding string in the title text for each term in the product thesaurus using an exact match. When a term in the product thesaurus matches the title text, that term constitutes the corresponding product term in the title text.
[0068] A third exemplary approach involves using semantic matching rules. After segmenting the title text, the semantic vectors of each segment are compared with the semantic vectors of each term in the product thesaurus. The term with the highest similarity exceeding a preset threshold is selected as the corresponding product term. When calculating the similarity, any data distance algorithm can be used, including but not limited to cosine similarity, Euclidean distance, Pearson correlation coefficient, Jaccard coefficient, etc. The product thesaurus can be pre-extracted by those skilled in the art from a large number of pre-collected product title texts to form the prior knowledge used in this application to determine the product terms in the title text.
[0069] Step S1220: Match the title text with a preset attribute dictionary to obtain knowledge entries that belong to attribute words in the title text;
[0070] One exemplary approach is to use rule-based matching to obtain multiple knowledge terms by segmenting the title text, then accurately match each knowledge term with a preset attribute term library, and identify the knowledge terms that match the attribute term library as attribute terms.
[0071] Another exemplary approach is to use rule-based matching to find the corresponding string in the title text for each term in the attribute dictionary using an exact match. When a term in the attribute dictionary matches the title text, that term constitutes the corresponding attribute word of the title text.
[0072] A third exemplary approach is to use semantic matching rules. After segmenting the title text, the semantic vectors of each segment are compared with the semantic vectors of each term in the attribute lexicon. The term with the highest similarity exceeding a preset threshold is selected as the corresponding attribute term. When calculating the similarity, any data distance algorithm can be used, including but not limited to cosine similarity, Euclidean distance, Pearson correlation coefficient, Jaccard coefficient, etc.
[0073] The attribute vocabulary library mentioned above can be pre-extracted by those skilled in the art from the title texts of a large number of pre-collected products, so as to constitute the prior knowledge used in this application to determine the attribute words of the title text.
[0074] Step S1230: Determine the word frequency characteristics of each knowledge term by referring to the statistical word frequency calculated from the preset title library;
[0075] The word frequency characteristics of each knowledge term in the title text can be determined by referring to prior knowledge provided by a pre-set title library. Specifically, at least one title library is prepared, which contains a large number of product title texts. Further, based on word segmentation of each title text in the title library, the statistical word frequency corresponding to each segment is calculated, and then the mapping relationship data between the segment and its word frequency is constructed into a word frequency statistics table.
[0076] E-commerce platform stores all have a product category system to categorize and organize their vast array of goods, meaning each product has its corresponding category. This category system can be multi-layered, containing multiple classification levels, each level containing multiple specific categories. The e-commerce platform can provide a standardized template for constructing the category system, which merchants can then modify and define themselves.
[0077] In one embodiment, a first title library can be pre-collected from the title texts of products belonging to the same category as the product corresponding to the title text. Additionally, a second title library can be pre-collected from the title texts of products from the same store as the product corresponding to the title text. Further, based on word segmentation of the first title library, word frequency statistics are performed as described above to obtain its corresponding first word frequency table. Similarly, based on word segmentation of the second title library, word frequency statistics are performed as described above to obtain its corresponding second word frequency table.
[0078] Therefore, when determining the word frequency features corresponding to each knowledge term in the title text, the statistical word frequency of the corresponding word segment is retrieved from the first word frequency statistics table and normalized to a corresponding value according to a preset normalization method, thus completing the construction of one word frequency feature for each knowledge term in the title text. Additionally, the statistical word frequency of the corresponding word segment is retrieved from the second word frequency statistics table and normalized to a corresponding value according to a preset normalization method, thus completing the construction of another word frequency feature for each knowledge term in the title text. One exemplary normalization method is to perform a logarithmic transformation or max-min normalization on the statistical word frequency to obtain the corresponding word frequency feature.
[0079] As can be seen, the word frequency features of each knowledge entry can be represented by one or more. Specifically, different word frequency statistics tables can be determined by different types of title libraries, and the statistical word frequency of each knowledge entry in different title libraries can be obtained, so that multiple word frequency features can be determined for each knowledge entry.
[0080] Step S1240: Determine the positional features of each knowledge term based on its position in the title text;
[0081] The position of each knowledge term in the title text is different, including its absolute position and its relative position to a knowledge term belonging to a product term. Therefore, by applying a preset normalization method to quantify the position information of each knowledge term in the title text into a numerical value, the position information can include absolute position information, relative position information, or both, and the position characteristics of each knowledge term can be obtained.
[0082] The order in which each knowledge term appears in the title text, i.e., its absolute position information, is processed using the following formula in the corresponding exemplary normalization method:
[0083] F abs = 1 / log2(1+L) abs )
[0084] Among them, L abs F represents the numerical value corresponding to the order of the knowledge entry in the title text. abs This refers to the absolute positional features obtained by applying the formula for normalization.
[0085] For each knowledge term belonging to product terms and attribute terms, normalization can be performed according to the above principles to determine the absolute position features corresponding to each knowledge term.
[0086] For each knowledge entry belonging to the attribute term, its relative position feature can be determined by referring to the relative position information of its closest product term. First, determine the difference in their ranking positions, and then normalize based on this difference. For example, in the title text "Fashionable and Elegant Shawl Cape Knitted Wool Sweater," the attribute term "fashionable" (position number 1) and its closest product term "shawl" (position number 5) have a relative position information of 4 obtained by calculating the difference in their position numbers. After determining this relative position information, the following formula can be applied to normalize it to obtain the corresponding relative position feature of the attribute term:
[0087] F rel = 1 / log2(1+L) rel )
[0088] Where: L rel F represents the numerical value corresponding to the relative position information of the attribute words. rel This refers to the relative position features obtained after normalization using the formula.
[0089] For each knowledge entry belonging to the product term, a standard value, such as the numerical value 1, can be used to describe the relative position information of the product term. Then, similarly to the attribute term, the normalization formula described above is applied to normalize the standard value to obtain the relative position feature corresponding to the product term.
[0090] Step S1250: Quantify and determine the information score of each knowledge term based on its frequency and positional features.
[0091] As mentioned earlier, the word frequency and positional features of each knowledge term in the title text have been determined and can be normalized to numerical features. Therefore, for each knowledge term, the information score corresponding to that knowledge term can be obtained by adding or weighting its word frequency and positional features. Thus, each product term and attribute term in the title text can obtain its corresponding information score. It is easy to understand that this information score integrates the positional and word frequency information value of the knowledge term, effectively measuring its contribution to the information value of the title text.
[0092] In this embodiment, each knowledge term in the title text and its corresponding word frequency features are determined by referring to prior knowledge, and the corresponding position features of each knowledge term in the title text are determined by referring to the position information of each knowledge term in the title text. The information value corresponding to each knowledge term is determined by quantifying the word frequency features and position features of each knowledge term, thereby realizing the effective quantification of the information value of the knowledge terms in the title text.
[0093] Please see Figure 3 In a further embodiment, step S1300, which involves calculating the similarity between multiple long texts composed of corresponding combinations of partial words in the title text and the title text, and selecting the long texts with higher similarity to construct a second candidate summary set, includes the following steps:
[0094] Step S1310: Obtain multiple long texts composed of corresponding combinations of some word elements in the title text;
[0095] The title text can be segmented to obtain multiple word units arranged in the order of the title text. An adjacent combination method can be used, which involves combining multiple word units that are adjacent in position within the title text. The specific number of word units combined can be determined by those skilled in the art as needed.
[0096] In one embodiment, to obtain multiple long texts composed of combinations of multiple words in adjacent positions in the title text as a summary of the title text, N-Gram segmentation can be used. Specifically, N-Gram segmentation with N>=2 is used. Those skilled in the art should know that N represents the number of words contained in the word extraction box. When N>=2, the word extraction box contains N words. Then, according to the preset movement step size for the word extraction box, the word extraction box is moved step by step to extract the N words in the word extraction box from the title text. Thus, it is realized that each step extracts a long text composed of N words in adjacent positions. After stepping, multiple long texts can be obtained. The movement step size can be set as needed, and it is recommended to set it to 1 to obtain a larger number of long texts.
[0097] Step S1320: Using a pre-trained text similarity model that has converged, calculate the similarity between each long text and the title text based on the semantic features of the title text and each of the long texts.
[0098] To filter the multiple long texts obtained and select those that are semantically highly similar to the title text, a pre-trained, converged text similarity model can be used to extract the semantic features corresponding to the title text and the multiple long texts, generating corresponding vectorized representations. Then, a similarity function is used to calculate the similarity between the title text and the vectorized representations of each long text. It can be understood that the similarity represents the degree of semantic similarity between the title text and each long text; a higher similarity indicates that the long text and the title text are more semantically similar.
[0099] The text similarity model can be a deep semantic learning-based network model in the field of NLP (Natural Language Processing) suitable for extracting semantic features from text. Specifically, it can adopt the open-source framework SentenceTransformers, which provides a large number of pre-trained Transformer models that have converged, such as BERT, RoBERTa, XLM-RoBERTa, and MPNet. The similarity function can be a cosine similarity function, Euclidean distance function, Pearson correlation coefficient function, Minkowski distance function, Mahalanobis distance function, Jaccard coefficient function, etc. Those skilled in the art can choose any one of them, as long as the vector distance between the title text and the corresponding vectorized representations of each long text can be calculated as the similarity.
[0100] Step S1330: Select long texts with similarity higher than a preset threshold to construct a second candidate summary set.
[0101] It is easy to understand that the long text is equivalent to extracting a portion of the title text. If it is highly similar to the title text semantically, it indicates that the extracted portion is a key point of the title text, meaning that the long text can serve as a summary of the title text to some extent. Accordingly, a second set of candidate summaries can be constructed by selecting long texts with a similarity exceeding a preset threshold. The specific value of the threshold can be set as needed by those skilled in the art.
[0102] In this embodiment, multiple long texts are constructed by extracting some words from the title text, and then the similarity between each long text and the title text is calculated. The long texts with higher similarity are selected to construct a second set of candidate summaries, thus scientifically and effectively generating long texts that can be used as summaries to a certain extent as candidate summaries of the title text.
[0103] In a further embodiment, in step S1310, the step of obtaining multiple long texts composed of corresponding combinations of partial words in the title text:
[0104] Step S1301: Segment the title text into words, and combine the resulting word units into adjacent combinations to obtain multiple corresponding long texts. The adjacent combination is the combination of multiple word units that are adjacent in position in the title text.
[0105] When segmenting the title text, for Chinese title text, algorithms such as jieba, Stanford, Hanlp, KCWS segmenter, THULAC, N-Gram, and deep learning can be used; for English title text, algorithms such as Keras, Spacy, Gensim, NNTK, N-Gram, and deep learning can be used.
[0106] As can be seen, after segmenting the title text, a corresponding segmented text can be obtained, which contains multiple word units arranged in the order of the title text.
[0107] It is understood that in order to obtain a long text that may serve as a summary of the title text, some words in the title text can be combined to obtain multiple long texts. Multiple words arranged adjacently in the word segmentation text can be combined by using an adjacent combination method.
[0108] In this embodiment, by combining adjacent words obtained from word segmentation of the title text, long texts that may serve as summaries of the title text can be generated easily and quickly, laying the foundation for selecting high-quality summaries from them in the future.
[0109] Please see Figure 4 In a further embodiment, step S1400, the training process of the text classification model, includes the following steps:
[0110] Step S1410: Obtain a single training sample from the prepared training set. Each training sample in the training set contains the title text of the product, a candidate summary, and a quality label. The quality label of the training sample indicates whether the candidate summary of the training sample can be used as a summary of the title text.
[0111] The system can collect sufficient title texts of multiple products corresponding to each category in the product category system. Then, referring to steps S1200-1300, a first candidate summary set and a second candidate summary set corresponding to each title text contained in each category can be constructed. Further, based on word segmentation of each candidate summary in the candidate summary set of each category, the word frequency corresponding to the last word segment in each candidate summary and its corresponding position in other candidate summaries is statistically analyzed. It can be understood that the word frequency corresponding to the product word segment is higher. However, the summary is usually a phrase or sentence composed of "several attribute words modifying the product word and the product word", and the product word is the word element placed at the end of the summary. The word element at the end can be the last word element in the summary or the word element immediately adjacent to it.
[0112] Therefore, for each of the last word segments, the word frequency corresponding to the last position of the candidate summary is greater than the minimum word frequency, for example, greater than a predetermined specific value and greater than or equal to the total number of title texts of the corresponding category multiplied by a predetermined proportion. The predetermined specific value and predetermined proportion can be set by those skilled in the art as needed. At the same time, the word frequency corresponding to the last position of the candidate summary is greater than the word frequency corresponding to its position in the candidate summary except for the position immediately adjacent to the last word. In addition, the word frequency corresponding to the last position of the candidate summary is greater than the word frequency of its position immediately adjacent to the last word in the candidate summary multiplied by a predetermined weight. The predetermined weight can be set by those skilled in the art as needed. This word segment is basically a product word.
[0113] Otherwise, the frequency of the word appearing at the last position in the candidate summary is less than the minimum frequency, for example, less than a predetermined specific value, which can be set as needed by those skilled in the art; or the frequency of the word appearing at the last position in the candidate summary is less than the frequency of the word appearing at each of the other positions in the candidate summary except the one immediately adjacent to the last position; or the frequency of the word appearing at the last position in the candidate summary is less than the frequency of the word appearing at each of the other positions in the candidate summary multiplied by a predetermined weight, which can be set as needed by those skilled in the art. This word is basically not a product word. Based on this, one or more word segments that are product words and one or more word segments that are not product words can be determined for each category.
[0114] For each candidate summary in the candidate summary set for each category, if the segmentation at the end of the candidate summary is a product term from the corresponding category, the candidate summary is labeled as a positive sample with a quality label of, for example, 1. If the segmentation at the end of the candidate summary is a non-product term from the corresponding category, the candidate summary is labeled as a negative sample with a quality label of, for example, 0. Each labeled candidate summary in the first and second candidate summary sets for each category is associated with its corresponding title text and quality label as a training sample. The title text and candidate summary in each training sample form a sentence pair as a data pair, thus constructing a training set.
[0115] The training set can be prepared by referring to the above disclosure and stored for retrieval and use in this step.
[0116] Step S1420: After the text classification model extracts the text semantic features of the training samples, the prediction module outputs the quality score corresponding to the prediction that the training samples are positive samples.
[0117] The text classification model encodes, calculates, and extracts bidirectional feature representations of the title text and candidate abstracts in the training samples to obtain corresponding vectorized representations. Then, the fully connected layer of the model, i.e. the prediction module, performs a linear transformation on the vectorized representations and maps them to a preset classification space. The classification space includes a positive class space and a negative class space. The positive class space represents training samples as positive samples, and the negative class space represents training samples as negative samples. The probability of mapping to the positive class space is then obtained as the quality score.
[0118] The text classification model can be a Bert series model, a GPT series model, XLNet, Bart, T5, etc., and those skilled in the art can select one of them as the text classification model as needed.
[0119] Step S1430: Calculate the loss value of the quality score of the text classification model based on the quality label corresponding to the training sample. If the loss value of the model does not reach the preset threshold, update the weights of the model and continue to call other training samples to carry out iterative training until the model converges.
[0120] The preset cross-entropy loss function is invoked. This function can be flexibly set by those skilled in the art based on prior knowledge or experimental experience. The cross-entropy loss value of the quality score of the text classification model is calculated based on the quality labels corresponding to the training samples. When the loss value reaches a preset threshold, it indicates that the model has been trained to a convergent state, and the model training can be terminated. When the loss value does not reach the preset threshold, it indicates that the model has not converged. Therefore, the model is updated with gradients based on the loss value. Usually, the weight parameters of each part of the model are corrected through backpropagation to make the model closer to convergence. Then, the next sample data in the training set is called to iteratively train the model until the model is trained to a convergent state.
[0121] In this embodiment, training samples labeled with quality tags are used to supervise the training of the text classification model. After the model is trained to convergence, it can predict the sentence pairs consisting of the input title text and the summary, and accurately predict the quality score that represents the summary as the summary of the title text.
[0122] Please see Figure 5 In a further embodiment, before step S1410, the step of obtaining a single training sample in the prepared training set, the following step is also included:
[0123] Step S1401: Obtain the title text of multiple products corresponding to each category in the product category system, and construct the corresponding first candidate summary set and second candidate summary set;
[0124] E-commerce platform stores all have a product category system in place to categorize and organize their vast array of goods. This category system can be multi-layered, containing multiple classification levels, each level encompassing multiple specific product categories. The e-commerce platform can provide a standardized template for constructing the category system, which merchants can then modify and finalize themselves.
[0125] You can refer to steps S1200-1300 to construct the first candidate summary set and the second candidate summary set corresponding to each title text included in each category.
[0126] Step S1402: Segment each candidate summary in the first candidate summary set and the second candidate summary set corresponding to each title text contained in each category. Construct a bag of words for each category using the bag-of-words model, which contains the word frequencies of each word segment corresponding to each category when it is in different positions in the candidate summary. Select the last word segment of each candidate summary of each category and associate it with the word frequencies of each position as an association pair to construct an association dataset.
[0127] For each category, each candidate summary in the first candidate summary set and the second candidate summary set corresponding to each title text is segmented. For Chinese candidate summaries, algorithms such as jieba, Stanford, Hanlp, KCWS segmenter, THULAC, N-Gram, and deep learning can be used. For English candidate summaries, algorithms such as Keras, Spacy, Gensim, NNTK, N-Gram, and deep learning can be used to obtain multiple segmented words corresponding to each candidate summary.
[0128] Each segment corresponding to each candidate summary is assigned a position number, and the position number is reversed according to the word order of each segment in the candidate summary. For example, in the candidate summary "Fitness WomenShorts", the position number corresponding to the segment "Fitness" is "3", the position number corresponding to the segment "Women" is "2", and the position number corresponding to the segment "Shorts" is "1".
[0129] Furthermore, a bag-of-words model is used to construct a bag-of-words for each category. This bag contains the frequency of each word segment corresponding to a particular category at different positions in the candidate summary. For example, candidate summary 1 is "Fitness Women Shorts", candidate summary 2 is "Comfortable Women Shorts", and candidate summary 3 is "TopsShorts Set". When the position number corresponding to the word "Shorts" is "1", the corresponding position frequency is "2"; when the position number corresponding to the word "Shorts" is "2", the corresponding position frequency is "1". The last word of each candidate summary for each category, i.e., the word corresponding to position number "1", is selected, and its corresponding position frequencies are associated as association pairs to construct an association dataset, thereby obtaining the association dataset corresponding to each word of each category.
[0130] Step S1403: Determine whether the frequency of all position words in each association pair in the association dataset satisfies all the conditions in the preset positive sample conditional distribution. If satisfied, label the corresponding candidate summary as the quality label corresponding to the positive sample. Otherwise, determine whether any one or more conditions in the preset negative sample conditional distribution are satisfied. If satisfied, label the corresponding candidate summary as the quality label corresponding to the negative sample.
[0131] In one embodiment, a positive sample conditional distribution and a negative sample conditional distribution can be preset. It is determined whether the frequency of all positional terms in each association pair in the association dataset satisfies all the conditions of the preset positive sample conditional distribution. If satisfied, it indicates that the word segment corresponding to that association pair belongs to a product term. Based on this, one or more product term segments and one or more non-product term segments corresponding to each product category can be determined. For each candidate summary in the first and second candidate summary sets for each product category, when the last word segment of the corresponding candidate summary is a product term of the corresponding product category, the last word segment can be the last word segment in the candidate summary or its immediately adjacent word segment. That is, when the last word segment or its immediately adjacent word segment is the previously determined product term segment, it indicates that the candidate summary meets the necessary conditions to be a summary of the corresponding title text, and the candidate summary is labeled with the quality tag corresponding to the positive sample. Otherwise, it is determined whether the frequency of all positional words in each association pair in the association dataset satisfies any one or more of the preset negative sample conditional distribution conditions. If they are satisfied, it indicates that the word segment corresponding to the association pair does not belong to product words. Based on this, one or more word segments that do not belong to product words corresponding to each category can be determined. For each candidate summary in the first and second candidate summary sets of each category, when the word segment at the end of the corresponding candidate summary is a word segment that does not belong to product words of the corresponding category, the word segment at the end can be the last word segment in the candidate summary or the word segment immediately adjacent to it. That is, when the last word segment or the word segment immediately adjacent to it is a word segment that is not a product word as previously determined, it indicates that the candidate summary does not meet the necessary conditions for being a summary of the corresponding title text, and the candidate summary is labeled with the quality label corresponding to the negative sample.
[0132] The positive sample conditional distribution includes: in the corresponding candidate summary, the positional word frequency at the last position of the candidate summary, i.e., the positional word frequency corresponding to the positional code "1" in the association pair, is greater than the positional word frequency at the immediately adjacent position, i.e., the positional word frequency corresponding to the positional code "2" in the association pair, multiplied by a first weight. The recommended first weight is 0.8, but those skilled in the art can also set it according to actual needs, for example, (positional word frequency corresponding to the positional code "1" in the association pair) > 0.8 * (positional word frequency corresponding to the positional code "2" in the association pair);
[0133] The frequency of the last position in the candidate summary is greater than the frequency of the position outside the immediately adjacent position. It is recommended to set corresponding weights to weight the greater-than relationship. The weights can be set according to actual needs, for example, (frequency of the position corresponding to "1" in the association pair) > 1.25 * (frequency of the position corresponding to "3" in the association pair), (frequency of the position corresponding to "1" in the association pair) > 1.25 * (frequency of the position corresponding to "4" in the association pair).
[0134] The positional frequency of the last position in the candidate summary, i.e., the positional frequency corresponding to the positional code "1" in the association pair, is greater than or equal to a first predetermined threshold. The first predetermined threshold is a specific value, such as 100. The specific value can be determined by those skilled in the art based on the total number of candidate summaries in the first and second candidate summary sets corresponding to the category of the candidate summary.
[0135] The positional frequency of the last position in the candidate summary, i.e. the positional frequency of the positional code "1" in the association pair, is greater than the second predetermined threshold, which is determined based on the total number of title texts contained in the category corresponding to the candidate summary.
[0136] If all conditions in the conditional distribution are met simultaneously, it can be determined that the positive sample conditional distribution is satisfied.
[0137] The negative sample conditional distribution includes: in the corresponding candidate summary, the positional word frequency of the last position in the candidate summary, that is, the positional word frequency corresponding to the position encoded as "1" in the association pair, is less than a first predetermined threshold;
[0138] The positional frequency of the last position in the candidate summary, which is the positional frequency of the position encoded as "1" in the association pair, is less than the positional frequency of the immediately adjacent position, which is the positional frequency of the position encoded as "2" in the association pair, multiplied by a second weight. The recommended second weight is 0.6, but those skilled in the art can also set it according to actual needs, for example, (positional frequency of the position encoded as "1" in the association pair) < 0.6 * (positional frequency of the position encoded as "2" in the association pair);
[0139] The frequency of the last position in the candidate summary is lower than the frequency of the position in other positions. For example, the frequency of the position corresponding to the position encoded as "1" in the association pair is less than the frequency of the position corresponding to the position encoded as "3" in the association pair, and the frequency of the position corresponding to the position encoded as "1" in the association pair is less than the frequency of the position corresponding to the position encoded as "4" in the association pair.
[0140] If any one or more of the conditions in the negative sample conditional distribution are met, it can be determined that the negative sample conditional distribution is met.
[0141] It is understandable that the above implementation of the annotation of candidate abstracts, on the one hand, accommodates the possibility that the determined product words are incomplete, and on the other hand, accommodates the characteristic that the context of the product title is not necessarily in complete word order, and there is a possibility that words representing the key semantics of the title text are adjacent to the product words.
[0142] Step S1404: Associate each labeled candidate abstract with its corresponding title text and quality label as a training sample to construct a training set.
[0143] Each labeled candidate summary in the candidate summary set of each category is associated with its corresponding title text and quality label as a training sample. The title text and candidate summary in each training sample form a sentence pair as a data pair to construct a training set.
[0144] In this embodiment, on the one hand, by statistically analyzing the positional word frequencies of the last word in each of the massive number of high-quality candidate summaries in the first and second candidate summary sets of each category, and determining whether the positional word frequencies corresponding to each last word satisfy the positive or negative sample conditional distribution, the massive number of high-quality candidate summaries can be labeled accordingly, clearly dividing a sufficient number of positive and negative samples. This makes the decision boundary of the text classification model trained with these positive and negative samples sufficiently smooth, improving the generalization ability, robustness, and accuracy of the text classification model. This lays the foundation for accurately predicting the quality score of the candidate summary corresponding to the title text in the actual application scenario using the model trained to convergence. On the other hand, it can save a lot of manpower costs required for labeling samples.
[0145] Please see Figure 6 This invention provides a product title summary generation device to meet one of the purposes of this application. It is a functional embodiment of the product title summary generation method of this application. The device includes a title acquisition module 1100, a first set construction module 1200, a second set construction module 1300, and a summary generation module 1400. Specifically: the title acquisition module 1100 is used to acquire the title text of the product; the first set construction module 1200 is used to extract knowledge terms belonging to product terms and attribute terms from the title text, determine the information score of each knowledge term based on its statistical characteristics, and select the corresponding product title summary based on the information score. The first candidate summary set is constructed by combining words and attribute words; the second set construction module 1300 is used to calculate the similarity between multiple long texts composed of corresponding combinations of some words in the title text and the title text, and select the long texts with higher similarity to construct the second candidate summary set; the summary generation module 1400 is used to form data pairs with the title text and each candidate summary in the first and second candidate summary sets, input them into a pre-trained and converged text classification model, predict the quality score corresponding to each candidate summary, and select the candidate summary with higher quality score as the summary of the title text.
[0146] In a further embodiment, the first set construction module 1200 includes: a first term matching submodule, used to match the title text with a preset product term library to obtain knowledge terms belonging to product terms in the title text; a second term matching submodule, used to match the title text with a preset attribute term library to obtain knowledge terms belonging to attribute terms in the title text; a term frequency feature submodule, used to determine the term frequency feature of each knowledge term by referring to the statistical term frequency calculated by the preset title library; a position feature submodule, used to determine the position feature of each knowledge term according to its position in the title text; and an information scoring submodule, used to quantify and determine the information score of each knowledge term based on its term frequency feature and position feature.
[0147] In a further embodiment, the second set construction module 1300 includes: a text acquisition submodule, used to acquire multiple long texts composed of corresponding combinations of some word units in the title text; a similarity calculation submodule, used to calculate the similarity between each long text and the title text based on the semantic features of the title text and each of the long texts using a pre-trained and converged text similarity model; and a text filtering submodule, used to filter out long texts with similarity higher than a preset threshold to construct a second candidate summary set.
[0148] In a further embodiment, the training process of the text classification model includes: a training set acquisition module, used to acquire a single training sample in a prepared training set, wherein each training sample in the training set contains the title text of a product, a candidate summary, and a quality label, and the quality label of the training sample indicates whether the candidate summary of the training sample can be used as a summary of the title text; a score prediction module, used to extract text semantic features from the training sample by the text classification model, and output the predicted quality score corresponding to the training sample as a positive sample through the prediction module; and an iterative training module, used to calculate the loss value of the quality score of the text classification model based on the quality label corresponding to the training sample, update the weights of the model when the model loss value does not reach a preset threshold, and continue to call other training samples to perform iterative training until the model converges.
[0149] In a further embodiment, before the training set acquisition module, the system further includes: a first set and a second set construction module 1300, used to acquire the title texts of multiple products corresponding to each category in the product category system, and construct corresponding first candidate summary sets and second candidate summary sets; a third set construction module, used to segment each candidate summary in the first candidate summary set and the second candidate summary set corresponding to each title text contained in each category, construct a bag of words for each category using a bag-of-words model, which includes the word frequencies of multiple positions corresponding to each word segment in the candidate summary when each word segment of the corresponding category is in a different position, select the last word of each candidate summary of each category, and associate it with its corresponding word frequencies as an association pair to construct an association dataset; a sample labeling module, used to determine whether the word frequencies of all positions in each association pair in the association dataset meet the preset condition distribution, if they meet the condition, label the corresponding candidate summary as the quality label corresponding to the positive sample, otherwise label it as the quality label corresponding to the negative sample; and a training set construction module, used to associate each candidate summary with its corresponding title text and quality label as training samples to construct a training set.
[0150] To address the aforementioned technical problems, embodiments of this application also provide computer equipment. For example... Figure 7 The diagram shows the internal structure of a computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected via a system bus. The computer-readable storage medium stores an operating system, a database, and computer-readable instructions. The database may store a sequence of control information. When the computer-readable instructions are executed by the processor, the processor can implement a product title summary generation method. The processor of the computer device provides computing and control capabilities to support the operation of the entire computer device. The memory of the computer device may store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute the product title summary generation method of this application. The network interface of the computer device is used for communication with a terminal. Those skilled in the art will understand that… Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0151] In this embodiment, the processor is used to execute... Figure 6The specific functions of each module and its sub-modules are defined within the device. The memory stores the program code and various data required to execute these modules or sub-modules. The network interface is used for data transmission between the user terminal and the server. In this embodiment, the memory stores the program code and data required to execute all modules / sub-modules in the product title summary generation device of this application. The server can call the server's program code and data to execute the functions of all sub-modules.
[0152] This application also provides a storage medium storing computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps of the product title summary generation method of any embodiment of this application.
[0153] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. This computer program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0154] In summary, on the one hand, this application employs two implementation methods to deeply mine the title text, initially generating a sufficient number of candidate abstracts of considerable quality. On this basis, a text classification model is used to accurately predict the quality score corresponding to each candidate abstract. Accordingly, high-quality candidate abstracts are selected as the abstracts of the title text.
[0155] On the other hand, this application determines the mathematical features of the last word of the summary corresponding to the training sample by statistics, matches it with a preset quantitative conditional distribution, and labels the corresponding training sample as a positive sample or a negative sample based on whether the mathematical features meet the conditional distribution. This eliminates the need for manual labeling of massive training samples, saving resource costs. In addition, the clear distinction between positive and negative samples ensures the performance of the text classification model trained on it, and can accurately predict the quality score corresponding to the summary in practical application scenarios.
[0156] Those skilled in the art will understand that the steps, measures, and solutions in the various operations, methods, and processes discussed in this application can be alternated, modified, combined, or deleted. Furthermore, other steps, measures, and solutions in the various operations, methods, and processes discussed in this application can also be alternated, modified, rearranged, decomposed, combined, or deleted. Furthermore, steps, measures, and solutions in the prior art that are similar to those disclosed in this application can also be alternated, modified, rearranged, decomposed, combined, or deleted.
[0157] The above description is only a partial embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for generating product title summaries, characterized in that, Includes the following steps: Get the product title text; Knowledge terms belonging to product terms and attribute terms are extracted from the title text. The information score of each knowledge term is determined by the statistical characteristics of the knowledge terms. Based on the information score, the corresponding combination text of product terms and attribute terms is selected to construct the first candidate summary set. Calculate the similarity between multiple long texts composed of partial word units from the title text and the title text, and select the long texts with higher similarity to construct a second candidate summary set; The title text is paired with each candidate summary in the first and second candidate summary sets to form a data pair. This pair is then input into a pre-trained and converged text classification model to predict the quality score corresponding to each candidate summary. The candidate summary with the higher quality score is selected as the summary of the title text. The training process of the text classification model includes: obtaining a single training sample from a pre-prepared training set, wherein each training sample in the training set contains the title text of a product, a candidate summary, and a quality label, and the quality label of the training sample indicates whether the candidate summary of the training sample can be used as a summary of the title text; after the text classification model extracts the text semantic features of the training sample, the prediction module outputs the quality score corresponding to the training sample being a positive sample; calculating the loss value of the quality score of the text classification model based on the quality label corresponding to the training sample; when the loss value of the model does not reach a preset threshold, the model weights are updated, and iterative training is continued using other training samples until the model converges; Before obtaining a single training sample from the prepared training set, the method further includes: obtaining the title texts of multiple products corresponding to each category in the product category system, constructing a corresponding first candidate summary set and a second candidate summary set; segmenting each candidate summary in the first and second candidate summary sets corresponding to each title text included in each category, constructing a bag-of-words model for each category, which includes the word frequencies of each word segment corresponding to different positions in the candidate summary, selecting the last word of each candidate summary of each category, and associating it with its corresponding word frequencies as association pairs to construct an association dataset; determining whether the word frequencies of all positions in each association pair in the association dataset satisfy all conditions in the preset positive sample conditional distribution, if satisfied, labeling the corresponding candidate summary as the quality label corresponding to the positive sample, otherwise determining whether it satisfies any one or more conditions in the preset negative sample conditional distribution, if satisfied, labeling the corresponding candidate summary as the quality label corresponding to the negative sample; and associating each labeled candidate summary with its corresponding title text and quality label as training samples to construct a training set. The positive sample conditional distribution includes: in the corresponding candidate summary, the positional word frequency of the last position in the candidate summary is greater than the positional word frequency of its immediate neighbor multiplied by a first weight; the positional word frequency of the last position in the candidate summary is greater than the positional word frequency of any position other than the immediate neighbor; the positional word frequency of the last position in the candidate summary is greater than or equal to a first predetermined threshold; and the positional word frequency of the last position in the candidate summary is greater than a second predetermined threshold, the second predetermined threshold being determined based on the total number of title texts contained in the category corresponding to the candidate summary. The negative sample conditional distribution includes: in the corresponding candidate summary, the frequency of the last position of the candidate summary is less than a first predetermined threshold; the frequency of the last position of the candidate summary is less than the frequency of the position of the immediately adjacent position multiplied by a second weight; and the frequency of the last position of the candidate summary is less than the frequency of the position of other positions.
2. The product title summary generation method according to claim 1, characterized in that, The step of extracting knowledge terms belonging to product terms and attribute terms from the title text, and determining the information score of each knowledge term based on its statistical characteristics, includes the following steps: The title text is matched with a preset product thesaurus to obtain the knowledge entries that belong to the product terms in the title text; The title text is matched with a preset attribute dictionary to obtain knowledge entries that belong to attribute words in the title text; The frequency characteristics of each knowledge term are determined by statistical word frequency calculation based on a pre-set title library; Determine the positional characteristics of each knowledge term based on its location within the title text; The information score of each knowledge term is determined by quantifying its word frequency and positional features.
3. The product title summary generation method according to claim 1, characterized in that, The step of calculating the similarity between multiple long texts composed of corresponding combinations of partial words in the title text and the title text, and selecting the long texts with higher similarity to construct a second candidate summary set, includes the following steps: Obtain multiple long texts composed of corresponding combinations of partial words from the title text; A pre-trained, convergent text similarity model is used to calculate the similarity between each long text and the title text based on the semantic features of the title text and the corresponding text of each long text. Long texts with similarity higher than a preset threshold are selected to construct a second set of candidate summaries.
4. The product title summary generation method according to claim 3, characterized in that, In the step of obtaining multiple long texts composed of corresponding combinations of partial words from the title text: The title text is segmented into words, and the resulting word units are combined into adjacent groups to obtain multiple corresponding long texts. The adjacent grouping is the combination of multiple word units that are adjacent in position in the title text.
5. A product title summary generation device, characterized in that, include: The title retrieval module is used to retrieve the title text of a product. The first set construction module is used to extract knowledge entries belonging to product words and attribute words from the title text, determine the information score of each knowledge entry based on the statistical characteristics of the knowledge entries, and select the corresponding combination text of product words and attribute words to construct the first candidate summary set based on the information score. The second set construction module is used to calculate the similarity between multiple long texts composed of corresponding combinations of partial words in the title text and the title text, and select the long texts with higher similarity to construct the second candidate summary set. The summary generation module is used to form data pairs with the title text and each candidate summary in the first candidate summary set and the second candidate summary set, input them into a pre-trained and converged text classification model, predict the quality score corresponding to each candidate summary, and select the candidate summary with the higher quality score as the summary of the title text. The training process of the text classification model includes: obtaining a single training sample from a pre-prepared training set, wherein each training sample in the training set contains the title text of a product, a candidate summary, and a quality label, and the quality label of the training sample indicates whether the candidate summary of the training sample can be used as a summary of the title text; after the text classification model extracts the text semantic features of the training sample, the prediction module outputs the quality score corresponding to the training sample being a positive sample; calculating the loss value of the quality score of the text classification model based on the quality label corresponding to the training sample; when the loss value of the model does not reach a preset threshold, the model weights are updated, and iterative training is continued using other training samples until the model converges; Before obtaining a single training sample from the prepared training set, the method further includes: obtaining the title texts of multiple products corresponding to each category in the product category system, constructing a corresponding first candidate summary set and a second candidate summary set; segmenting each candidate summary in the first and second candidate summary sets corresponding to each title text included in each category, constructing a bag-of-words model for each category, which includes the word frequencies of each word segment corresponding to different positions in the candidate summary, selecting the last word of each candidate summary of each category, and associating it with its corresponding word frequencies as association pairs to construct an association dataset; determining whether the word frequencies of all positions in each association pair in the association dataset satisfy all conditions in the preset positive sample conditional distribution, if satisfied, labeling the corresponding candidate summary as the quality label corresponding to the positive sample, otherwise determining whether it satisfies any one or more conditions in the preset negative sample conditional distribution, if satisfied, labeling the corresponding candidate summary as the quality label corresponding to the negative sample; and associating each labeled candidate summary with its corresponding title text and quality label as training samples to construct a training set. The positive sample conditional distribution includes: in the corresponding candidate summary, the positional word frequency of the last position in the candidate summary is greater than the positional word frequency of its immediate neighbor multiplied by a first weight; the positional word frequency of the last position in the candidate summary is greater than the positional word frequency of any position other than the immediate neighbor; the positional word frequency of the last position in the candidate summary is greater than or equal to a first predetermined threshold; and the positional word frequency of the last position in the candidate summary is greater than a second predetermined threshold, the second predetermined threshold being determined based on the total number of title texts contained in the category corresponding to the candidate summary. The negative sample conditional distribution includes: in the corresponding candidate summary, the frequency of the last position of the candidate summary is less than a first predetermined threshold; the frequency of the last position of the candidate summary is less than the frequency of the position of the immediately adjacent position multiplied by a second weight; and the frequency of the last position of the candidate summary is less than the frequency of the position of other positions.
6. The product title summary generation device according to claim 5, characterized in that, The first set construction module includes: a first term matching submodule, used to match the title text with a preset product term library to obtain knowledge terms belonging to product terms in the title text; a second term matching submodule, used to match the title text with a preset attribute term library to obtain knowledge terms belonging to attribute terms in the title text; a term frequency feature submodule, used to determine the term frequency feature of each knowledge term by referring to the statistical term frequency calculated by the preset title library; a position feature submodule, used to determine the position feature of each knowledge term according to its position in the title text; and an information scoring submodule, used to quantify and determine the information score of each knowledge term based on its term frequency feature and position feature.
7. The product title summary generation device according to claim 5, characterized in that, The second set construction module includes: a text acquisition submodule, used to acquire multiple long texts composed of corresponding combinations of some word units in the title text; a similarity calculation submodule, used to calculate the similarity between each long text and the title text based on the semantic features of the title text and each of the long texts using a pre-trained and converged text similarity model; and a text filtering submodule, used to filter out long texts with similarity higher than a preset threshold to construct a second candidate summary set.
8. The product title summary generation device according to claim 5, characterized in that, The training process of the text classification model includes: a training set acquisition module, used to acquire a single training sample in a prepared training set, wherein each training sample in the training set contains the title text of the product, a candidate summary, and a quality label, and the quality label of the training sample indicates whether the candidate summary of the training sample can be used as a summary of the title text; a score prediction module, used to extract the text semantic features of the training sample by the text classification model, and output the predicted quality score corresponding to the training sample being a positive sample; and an iterative training module, used to calculate the loss value of the quality score of the text classification model based on the quality label corresponding to the training sample, and to update the weights of the model when the model loss value does not reach a preset threshold, and to continue to call other training samples to perform iterative training until the model converges.
9. A computer device comprising a central processing unit and a memory, characterized in that, The central processing unit is used to invoke and run a computer program stored in the memory to perform the steps of the method as described in any one of claims 1 to 4.
10. A computer-readable storage medium, characterized in that, It stores, in the form of computer-readable instructions, a computer program implemented according to any one of claims 1 to 4, which, when invoked by a computer, executes the steps included in the corresponding method.
Citation Information
Patent Citations
Text abstract generation method and device, text abstract training method and device, equipment and medium
CN111026861A
Public opinion text abstract extraction method and device, equipment and computer storage medium
CN114201600A
Commodity title abstract generation method and device, equipment, medium and product
CN114780687A