Method and apparatus for outputting information

By extracting character vectors and position encoding vectors in item description text, and using reinforcement learning to train the selling point word recognition model, the problem of difficulty in mining high-quality selling point words in the existing technology is solved, and the effect of improving click-through rate and conversion rate is achieved.

CN113761878BActive Publication Date: 2025-05-23BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010626171.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-02
Publication Date
2025-05-23
Estimated Expiration
2040-07-02

AI Technical Summary

Technical Problem

The prior art is difficult to discover high-quality selling words that users are interested in, and it is easy to concentrate on high-frequency words, and high-frequency does not necessarily mean high conversion rates.

Method used

By obtaining the description text of the item, extracting the character vector and position encoding vector of characters, and inputting them into the pre-trained selling point word recognition model, the model is trained using reinforcement learning methods to optimize conversion rate and click-through rate-related indicators, thereby generating and outputting high-quality selling point words.

Benefits of technology

It realizes effective mining of high-quality selling words, improves the click-through rate and conversion rate of product information, and can more accurately identify selling words that users are interested in.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113761878B_ABST
    Figure CN113761878B_ABST
Patent Text Reader

Abstract

The present application embodiment discloses a method and device for outputting information. A specific implementation of the method includes: obtaining a description text of an item, extracting a character vector and a position encoding vector of a character in the description text, wherein the position encoding vector is used to characterize the position of the character in the description text; inputting the input vector into a pre-trained selling point word recognition model to obtain the label of the character in the description text and the probability of the character corresponding to the character, wherein the selling point word recognition model is based on the relevant indicators corresponding to the selling point words in a preset selling point word set, and is trained using a reinforcement learning method, and the relevant indicators include at least one of the following: conversion rate and click-through rate; based on the labels of the characters in the description text, a selling point word set is generated; based on the probability of the characters in the selling point words corresponding to the selling point words, the probability of the selling point words corresponding to the selling point words is determined, and based on the probability of the selling point words corresponding to the selling point words, a selling point word is selected from the selling point word set for output. This implementation can mine out higher-quality selling point words.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technology, and in particular to a method and device for outputting information. Background Art

[0002] Mining selling words for products refers to extracting valuable and eye-catching words from the input text. In the field of e-commerce, high-quality selling words can not only help users quickly understand the characteristics of products, but also improve product conversion rates and bring profits to merchants. Mining selling words for products is usually regarded as a sequence labeling problem. Traditional selling word mining based on sequence labeling strategies is affected by the cross-entropy optimization objective function, and the mined selling words are often concentrated on high-frequency words. However, high frequency does not mean high online conversion rate. Many low-frequency selling words with high conversion rates are difficult to mine with existing methods. How to mine high-quality selling words that users are interested in, rather than the same general high-frequency selling words, is of great significance to each e-commerce company. Summary of the invention

[0003] The embodiments of the present application provide a method and device for outputting information.

[0004] In a first aspect, an embodiment of the present application provides a method for outputting information, including: obtaining a description text of an item, extracting character vectors and position coding vectors of characters in the description text, wherein the position coding vector is used to characterize the position of the character in the description text; inputting the input vector into a pre-trained selling point word recognition model to obtain labels of characters in the description text and the probability of character correspondence, wherein the input vector is determined based on the character vector and the position coding vector of the character, the label includes a label used to characterize that the character is a character in a selling point word, and the selling point word recognition model is trained using a reinforcement learning method based on relevant indicators corresponding to selling point words in a preset selling point word set, and the relevant indicators include at least one of the following: conversion rate and click-through rate; generating a selling point word set based on the labels of the characters in the description text; determining the probability of selling point word correspondence based on the probability of character correspondence in the selling point word, and selecting selling point words from the selling point word set for output based on the probability of selling point word correspondence.

[0005] In a second aspect, an embodiment of the present application provides a device for outputting information, including: an acquisition unit, configured to acquire a description text of an item, extract a character vector and a position coding vector of a character in the description text, wherein the position coding vector is used to characterize the position of the character in the description text; an input unit, configured to input the input vector into a pre-trained selling point word recognition model, and obtain a label of the character in the description text and a probability of the character corresponding to the character, wherein the input vector is determined based on the character vector and the position coding vector of the character, the label includes a label for characterizing that the character is a character in a selling point word, and the selling point word recognition model is based on relevant indicators corresponding to selling point words in a preset selling point word set, and is trained using a reinforcement learning method, and the relevant indicators include at least one of the following: conversion rate and click-through rate; a generation unit, configured to generate a selling point word set based on the labels of the characters in the description text; an output unit, configured to determine the probability of the selling point word corresponding to the character in the selling point word, and select a selling point word from the selling point word set for output based on the probability of the selling point word corresponding to the character, and select a selling point word from the selling point word set for output based on the probability of the selling point word corresponding to the selling point word.

[0006] In a third aspect, an embodiment of the present application provides an electronic device, comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation method in the first aspect.

[0007] In a fourth aspect, an embodiment of the present application provides a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method described in any implementation manner in the first aspect.

[0008] The method and device for outputting information provided by the above-mentioned embodiment of the present application obtains the description text of the item, extracts the character vector and position encoding vector of the characters in the above-mentioned description text; then, inputs the input vector into the pre-trained selling point word recognition model to obtain the label of the characters in the above-mentioned description text and the probability corresponding to the characters, wherein the above-mentioned selling point word recognition model is based on the conversion rate and / or click-through rate corresponding to the selling point words in the preset selling point word set, and is trained by using a reinforcement learning method; then, based on the labels of the characters in the above-mentioned description text, a selling point word set is generated; finally, based on the probability corresponding to the characters in the selling point words, the probability corresponding to the selling point words is determined, and based on the probability corresponding to the selling point words, a selling point word is selected from the above-mentioned selling point word set for output. This method takes the online indicators of improving the click-through rate and / or conversion rate as the optimization goal, and performs reinforcement learning on the selling point word recognition model. In this way, more high-quality selling point words can be mined, thereby improving the click-through rate and / or conversion rate of commodity information. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Other features, objects and advantages of the present application will become more apparent by reading the detailed description of non-limiting embodiments made with reference to the following drawings:

[0010] Figure 1 is an exemplary system architecture diagram to which various embodiments of the present application may be applied;

[0011] Figure 2 is a flow chart of an embodiment of a method for outputting information according to the present application;

[0012] Figure 3 is a flow chart of another embodiment of a method for outputting information according to the present application;

[0013] Figure 4 It is a schematic diagram of training a selling point word recognition model using a reinforcement learning method according to the method for outputting information of the present application;

[0014] Figure 5 is a schematic structural diagram of an embodiment of a device for outputting information according to the present application;

[0015] Figure 6 It is a structural diagram of a computer system suitable for implementing an electronic device of an embodiment of the present application. DETAILED DESCRIPTION

[0016] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It should also be noted that, for ease of description, only the parts related to the relevant invention are shown in the accompanying drawings.

[0017] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0018] Figure 1 An exemplary system architecture 100 is shown to which an embodiment of the method for outputting information of the present application may be applied.

[0019] like Figure 1As shown, the system architecture 100 may include user terminals 1011, 1012, networks 1021, 1022, a server 103, and output terminals 1041, 1042, 1043. The network 1021 is used to provide a medium for a communication link between the user terminals 1011, 1012 and the server 103. The network 1022 is used to provide a medium for a communication link between the server 103 and the output terminals 1041, 1042, 1043. The networks 1021, 1022 may include various connection types, such as wired, wireless communication links, or optical fiber cables, etc.

[0020] The user terminals 1011 and 1012 can interact with the server 103 through the network 1021 to send or receive messages (for example, the server 103 can obtain the description text of the item from the user terminals 1011 and 1012). Various communication client applications can be installed on the user terminals 1011 and 1012, such as shopping applications, text editing applications, and instant messaging software.

[0021] User terminals 1011 and 1012 can be hardware or software. When user terminals 1011 and 1012 are hardware, they can be various electronic devices that support information interaction, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, etc. When user terminals 1011 and 1012 are software, they can be installed in the electronic devices listed above. They can be implemented as multiple software or software modules, or as a single software or software module. No specific limitation is made here.

[0022] The output terminals 1041, 1042, 1043 can interact with the server 103 through the network 1022 to send or receive messages, etc. (for example, the output terminals 1041, 1042, 1043 can receive the selling point words output by the server 103). Various communication client applications, such as shopping applications, instant messaging software, etc., can be installed on the output terminals 1041, 1042, 1043.

[0023] The output terminals 1041, 1042, 1043 may be hardware or software. When the output terminals 1041, 1042, 1043 are hardware, they may be various electronic devices that support information interaction, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, etc. When the output terminals 1041, 1042, 1043 are software, they may be installed in the electronic devices listed above. They may be implemented as multiple software or software modules, or as a single software or software module. No specific limitation is made here.

[0024] The server 103 may be a server that provides various services. For example, it may be a background server that analyzes the description text of an item. The server 103 may first obtain the description text of the item from the user terminals 1011 and 1012, and extract the character vectors and position coding vectors of the characters in the description text; then, the input vector determined based on the character vector and the position coding vector of the character may be input into a pre-trained selling point word recognition model to obtain the labels of the characters in the description text and the probability of the characters corresponding to the characters; then, a selling point word set may be generated based on the labels of the characters in the description text; finally, the probability of the selling point words corresponding may be determined based on the probability of the characters in the selling point words, and the selling point words may be selected from the selling point word set for output based on the probability of the selling point words corresponding to the characters, for example, the selected selling point words may be output to the output terminals 1041, 1042, and 1043.

[0025] It should be noted that the server 103 can be hardware or software. When the server 103 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or it can be implemented as a single server. When the server 103 is software, it can be implemented as multiple software or software modules (for example, for providing distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.

[0026] It should be noted that the method for outputting information provided in the embodiment of the present application is usually executed by the server 103.

[0027] It should be noted that the server 103 may store the description text of the item locally, and the server 103 may obtain the description text of the item locally. In this case, the exemplary system architecture 100 may not have the user terminals 1011 and 1012 and the network 1021.

[0028] It should also be noted that the server 103 may be connected to a display device (eg, a display screen) to display the output selling point words. In this case, the exemplary system architecture 100 may not have the network 1022 and the output terminals 1041 , 1042 , 1043 .

[0029] It should be understood that Figure 1 The number of user terminals, networks, servers and output terminals in the embodiment is only for illustration. Any number of user terminals, networks, servers and output terminals may be provided according to implementation requirements.

[0030] Continue to refer Figure 2 , shows a process 200 of an embodiment of a method for outputting information according to the present application. The method for outputting information comprises the following steps:

[0031] Step 201, obtain the description text of the item, and extract the character vector and position code vector of the characters in the description text.

[0032] In this embodiment, the execution subject of the method for outputting information (for example Figure 1 The server shown in the figure can obtain the description text of the item. Here, the above-mentioned items may include but are not limited to: goods and services. The description text of the above-mentioned items may include but are not limited to at least one of the following: the product title and the product details page text. Among them, the above-mentioned product details page text may be the text obtained by performing OCR (Optical Character Recognition) processing on the product details image.

[0033] Afterwards, the execution entity can extract the character vector and position encoding vector of the character in the description text. The position encoding vector can be used to characterize the position of the character in the description text. Here, vector extraction algorithms such as BoW model (Bag of words model), simhash algorithm and word2vec algorithm can be used to extract the character vector of the character from the description text. It should be noted that vector extraction algorithms such as BoW model, simhash algorithm and word2vec algorithm are currently well-known technologies that are widely studied and applied, and will not be repeated here.

[0034] Here, character vectors are used as model input to avoid the impact of errors in word segmentation tools on the annotation boundaries and reduce the occurrence of OOV (Out of vocabulary) problems.

[0035] Step 202: Input the input vector into a pre-trained selling point word recognition model to obtain labels describing characters in the text and the corresponding probabilities of the characters.

[0036] In this embodiment, the execution subject may input the input vector into a pre-trained selling point word recognition model to obtain the labels of the characters in the description text and the corresponding probabilities of the characters. Here, the input vector is usually determined based on the character vector and position encoding vector of the character. As an example, the character vector and position encoding vector of the character may be concatenated, and the concatenated result may be used as the input vector.

[0037] Here, selling points usually refer to the unprecedented, unique or distinctive features or characteristics of the goods being sold. Selling point words are words that describe the features or characteristics of the goods.

[0038] In this embodiment, the label of the character may include a label for indicating that the character is a character in the selling point word, and may also include a label for indicating that the character is not a character in the selling point word. Here, the label may adopt the marking method of {B, I, O}, that is, the character at the beginning of the selling point word is usually marked as B, the character in the selling point word is usually marked as I, and the other characters are usually marked as O. The probability corresponding to the character usually refers to the probability that the character is the determined label.

[0039] In this embodiment, the above-mentioned selling point word recognition model is generally used to characterize the corresponding relationship between the input vector corresponding to the text and the label of the character in the text and the probability corresponding to the character. The above-mentioned selling point word recognition model can be trained by using a reinforcement learning method based on the relevant indicators corresponding to the selling point words in the preset selling point word set. Here, the above-mentioned relevant indicators may include at least one of the following: conversion rate and click-through rate. Conversion rate (Conversion Rate, CVR) generally refers to the conversion rate of users clicking on advertisements to becoming an effectively activated or registered or even paying user, which is usually the ratio of conversion volume to click volume. Click-through rate (Click Through Rate, CTR) can also be called click-through rate or click-through rate, which is usually the ratio of the actual number of clicks on the product advertisement to the number of advertisements displayed.

[0040] In this embodiment, reinforcement learning (RL), also known as reinforcement learning, evaluation learning or enhanced learning, is used to describe and solve the problem of an intelligent agent (i.e., the above-mentioned selling point word recognition model) using learning strategies to maximize returns or achieve specific goals during the interaction with the environment. In standard reinforcement learning, the intelligent agent, as a learning system, obtains the current state information S of the external environment, takes a tentative action A to the environment, and obtains the reward value R and the new environmental state for this action from the environmental feedback. If an action A of the intelligent agent results in a positive reward (immediate reward), the tendency of the intelligent agent to perform this action in the future will be strengthened; otherwise, the tendency of the intelligent agent to perform this action will be weakened. In the repeated interaction between the control behavior of the learning system and the state and evaluation of the environmental feedback, the mapping strategy from state to action is continuously modified in a learning manner to achieve the purpose of optimizing system performance.

[0041] Here, the state information S = (skuid, sp, ctr, cvr), where skuid represents the unique code of the product, sp represents the selling point words mined for the product, ctr represents the click-through rate, and cvr represents the conversion rate. The action space A contains all the actions that the agent can take in each state. In this scenario, it is to select the candidate selling point word list of the current product (sp 1 ,sp 2 ,...,sp n) as the action of the current selling point word. According to the online indicators of e-commerce, the click-through rate and conversion rate indicators can be used as specific reward functions to guide the agent to mine more high-quality (high click-through rate and high conversion rate) selling point words during the learning process.

[0042] Step 203: Generate a set of selling point words based on the labels of the characters in the description text.

[0043] In this embodiment, the execution subject may generate a set of selling point words based on the labels of the characters in the description text obtained in step 202. Specifically, the execution subject may use the labels of the characters in the description text to splice adjacent characters in the characters representing the selling point words to obtain the selling point words. As an example, if the labels corresponding to the characters in the description text "New Summer Splashed Ink Graffiti Little White Shoes" are OOOOBIIIOOOO, the execution subject may splice the adjacent characters corresponding to BIII to obtain the selling point word "Splashed Ink Graffiti".

[0044] Step 204, based on the probability corresponding to the characters in the selling point words, determine the probability corresponding to the selling point words, and based on the probability corresponding to the selling point words, select a selling point word from the selling point word set for output.

[0045] In this embodiment, the execution entity may determine the probability corresponding to the selling point word based on the probability corresponding to the characters in the selling point word. Specifically, the execution entity may determine the product of the probabilities corresponding to the characters in the selling point word as the probability corresponding to the selling point word. The execution entity may also determine the average value of the probabilities corresponding to the characters in the selling point word as the probability corresponding to the selling point word.

[0046] Afterwards, the execution entity may select and output selling point words from the set of selling point words generated in step 203 based on the probabilities corresponding to the selling point words. As an example, the execution entity may select and output a preset number of selling point words from the set of selling point words in descending order of corresponding probabilities.

[0047] The method provided in the above-mentioned embodiment of the present application takes improving the online indicators of click-through rate and / or conversion rate as the optimization goal, and performs reinforcement learning on the selling point word recognition model. In this way, better selling point words can be discovered.

[0048] In some optional implementations of this embodiment, the strategy of the above selling point word recognition model may be determined by the following formula (1):

[0049]

[0050] Among them, γ represents the discount rate, γ∈[0,1], γ kUsually, the current reward is more important than the future reward. k represents the number of executions, t represents the current moment, and r represents the current time. t+k Indicates the reward value obtained by the selling point word recognition model when it is executed for the kth time, s t Represents the current state. Represents the cumulative reward value obtained after performing a set of actions at the current state. represents the expected value of the cumulative reward value, argmax(f(x)) is the variable point x (or set of x) corresponding to the maximum value of f(x), and π represents the selling point word selection path corresponding to the maximum value of the expected value of the cumulative reward value. The optimization goal of the agent (the above selling point word recognition model) is to find an optimal strategy π so that the maximum long-term cumulative reward value can be obtained in any state s and any time t.

[0051] As an example, if the current state S t =(skuid=1,sp=lace,ctr=0.15,cvr=0.03), after executing the next action (selecting sp=knitting pattern), ctr=0.62,cvr=0.12, and the corresponding reward value can be determined based on ctr and cvr. The agent can select actions with different selling point words and select the action path with the maximum reward as the optimal solution.

[0052] In some optional implementations of this embodiment, the above-mentioned reward value may be a weighted average of the conversion rate and click-through rate corresponding to the selling point words, and the weights corresponding to the conversion rate and click-through rate may be preset experience values. Here, the reward value may be determined by the following formula (2):

[0053] r=0.7×ctr+0.3×cvr (2)

[0054] Among them, r represents the reward value, ctr represents the click rate, and cvr represents the conversion rate.

[0055] As an example, if ctr=0.62, cvr=0.12, the corresponding reward value may be 0.7×0.62+0.3×0.12=0.47.

[0056] In some optional implementations of this embodiment, for each selling point word in the above-mentioned selling point word set, the above-mentioned execution subject may determine whether the selling point word corresponds to a relevant indicator. If there is no relevant indicator for the selling point word, the above-mentioned execution subject may obtain the k nearest neighbor selling point words of the selling point word. Here, the k nearest neighbor selling point words of the selling point word may be the k selling point words whose word vectors are closest to the word vector of the selling point word. The above-mentioned execution subject may determine the weighted average of the relevant indicators corresponding to the above-mentioned k nearest neighbor selling point words as the relevant indicator corresponding to the selling point word. Among them, for each neighbor selling point word in the above-mentioned k nearest neighbor selling point words, the weight of the neighbor selling point word may be the similarity between the neighbor selling point word and the selling point word. Here, the similarity between the neighbor selling point word and the selling point word may be the distance between the word vector of the neighbor selling point word and the word vector of the selling point word. The distance between word vectors may be a cosine distance or a Euclidean distance.

[0057] Here, the execution entity can determine the conversion rate of the selling point word through the following formula (3):

[0058]

[0059] Among them, cvr represents the conversion rate, i∈[1,k], cvr i represents the conversion rate corresponding to the i-th selling point word, sp represents the selling point word, sp i Represents the i-th selling point word, sim(sp i ,sp) represents the similarity between the selling point word and the i-th selling point word, and k represents the number of neighboring selling point words.

[0060] Here, the execution entity can determine the click rate of the selling point word by the following formula (4):

[0061]

[0062] Among them, ctr represents the click rate, i∈[1,k], ctr i represents the click-through rate corresponding to the i-th selling point word, sp represents the selling point word, sp i Represents the i-th selling point word, sim(sp i ,sp) represents the similarity between the selling point word and the i-th selling point word, and k represents the number of neighboring selling point words.

[0063] It should be noted that the word vectors of the selling point words and the neighboring selling point words can be determined based on the Word2vec model.

[0064] Further references Figure 3 , which shows a process 300 of another embodiment of a method for outputting information. The process 300 of the method for outputting information includes the following steps:

[0065] Step 301, obtain the description text of the item, and extract the character vector and position encoding vector of the characters in the description text.

[0066] In this embodiment, the specific operation of step 301 is already Figure 2 In the illustrated embodiment, step 201 is introduced in detail and will not be repeated here.

[0067] Step 302: input the input vector into the bidirectional encoding layer of the selling point word recognition model to obtain a first vector.

[0068] In this embodiment, the selling point word recognition model generally includes a bidirectional encoding layer, a bidirectional long short-term memory (Bi-directional Long Short-Term Memory, Bi-LSTM) layer and a conditional random field (Conditional Random Field, CRF) layer. Among them, the bidirectional long short-term memory layer is usually composed of a forward LSTM and a backward LSTM, both of which are often used to model context information in natural language processing tasks. Conditional random fields are a discriminative probability model that is often used to annotate or analyze sequence data, such as natural language text or biological sequences.

[0069] In this embodiment, the execution subject of the method for outputting information (for example Figure 1 The server shown in the figure can input the input vector into the bidirectional encoding layer of the above-mentioned selling point word recognition model to obtain a first vector. Here, the above-mentioned first vector is usually a character vector that incorporates the contextual semantic relationship. The above-mentioned input vector is usually determined based on the character vector and the position encoding vector of the character. As an example, the character vector and the position encoding vector of the character can be spliced, and the splicing result can be used as the input vector.

[0070] As an example, the above bidirectional encoding layer can be a BERT (Bidirectional Encoder Representation from Transformers) model. The BERT model uses Masked LM (Masked Language Model) and Next Sentence Prediction to capture word and sentence level expressions respectively. Masked LM randomly masks part of the input words and then predicts those masked words. The purpose of Next Sentence Prediction is to enable the model to understand the connection between two sentences.

[0071] Step 303: input the first vector into the bidirectional long short-term memory layer of the selling point word recognition model to obtain a second vector.

[0072] In this embodiment, the execution subject may input the first vector obtained in step 302 into the bidirectional long short-term memory layer of the selling point word recognition model to obtain a second vector. Here, the second vector is usually a character vector that incorporates contextual semantic relations again based on the first vector.

[0073] Step 304: input the second vector into the conditional random field layer of the selling point word recognition model to obtain labels describing characters in the text and the corresponding probabilities of the characters.

[0074] In this embodiment, the execution subject may input the second vector obtained in step 303 into the conditional random field layer of the selling point word recognition model to obtain the labels of the characters in the description text and the probability of the characters corresponding to the labels. The labels of the characters may include labels used to characterize that the characters are characters in the selling point words, and may also include labels used to characterize that the characters are not characters in the selling point words. Here, the labels may be marked in the form of {B, I, O}, that is, the characters at the beginning of the selling point words are usually marked as B, the characters in the selling point words are usually marked as I, and the other characters are usually marked as O. The probability of a character corresponding to a character usually refers to the probability that the character is the determined label.

[0075] In this embodiment, the above-mentioned selling point word recognition model can be trained by using a reinforcement learning method based on relevant indicators corresponding to the selling point words in a preset selling point word set. Here, the above-mentioned relevant indicators can include at least one of the following: conversion rate and click-through rate. Conversion rate usually refers to the conversion rate from a user clicking on an advertisement to becoming an effectively activated or registered or even a paying user, which is usually the ratio of the conversion volume to the click volume. Click-through rate can also be called click-through rate or click-through rate, which is usually the ratio of the actual number of clicks on a product advertisement to the number of advertisements displayed.

[0076] In this embodiment, reinforcement learning, also known as reinforcement learning, evaluation learning or enhanced learning, is used to describe and solve the problem of an intelligent agent (i.e., the above-mentioned selling point word recognition model) using learning strategies to maximize rewards or achieve specific goals during the interaction with the environment. In standard reinforcement learning, the intelligent agent, as a learning system, obtains the current state information S of the external environment, takes a tentative action A to the environment, and obtains the reward value R and the new environmental state for this action from the environmental feedback. If an action A of the intelligent agent results in a positive reward, the tendency of the intelligent agent to perform this action in the future will be strengthened; otherwise, the tendency of the intelligent agent to perform this action will be weakened. In the repeated interaction between the control behavior of the learning system and the state and evaluation of the environmental feedback, the mapping strategy from state to action is continuously modified in a learning manner to achieve the purpose of optimizing system performance.

[0077] Here, the state information S = (skuid, sp, ctr, cvr), where skuid represents the unique code of the product, sp represents the selling point words mined for the product, ctr represents the click-through rate, and cvr represents the conversion rate. The action space A contains all the actions that the agent can take in each state. In this scenario, it is to select the candidate selling point word list of the current product (sp 1 ,sp 2 ,...,sp n ) as the action of the current selling point word. According to the online indicators of e-commerce, the click-through rate and conversion rate indicators can be used as specific reward functions to guide the agent to mine more high-quality (high click-through rate and high conversion rate) selling point words during the learning process.

[0078] Step 305: Generate a set of selling point words based on the labels of the characters in the description text.

[0079] Step 306, determining the probability of the selling point word corresponding to the characters in the selling point word based on the probability of the characters in the selling point word corresponding to the characters, and selecting a selling point word from the selling point word set for output based on the probability of the selling point word corresponding to the characters.

[0080] In this embodiment, the specific operations of steps 305-306 are already Figure 2 In the illustrated embodiment, steps 203 - 204 are described in detail and will not be repeated here.

[0081] from Figure 3 It can be seen that Figure 2 Compared with the corresponding embodiment, the process 300 of the method for outputting information in this embodiment embodies the steps of inputting the input vector into the bidirectional encoding layer of the selling point word recognition model, inputting the output result of the bidirectional encoding layer into the bidirectional long short-term memory layer of the selling point word recognition model, and inputting the output result of the bidirectional long short-term memory layer into the conditional random field layer of the selling point word recognition model to obtain the label describing the characters in the text and the probability of the characters corresponding. Therefore, the scheme described in this embodiment can make the character vector incorporate more contextual semantic relationships, so as to obtain more accurate character labels and the probability of character corresponding.

[0082] Continue to see Figure 4 , Figure 4 FIG. 1 is a schematic diagram of a method for outputting information according to this embodiment of the present invention, using a reinforcement learning method to train a selling point word recognition model. Figure 4In the schematic diagram, the selling point word recognition model includes an Embeding layer, a Bert layer, a Bi-LSTM layer, and a CRF layer. The execution subject can input the description text 401 "New small fresh lace-up black casual shoes" into the Embeding layer to obtain the character vector and position encoding vector of each character; then, the execution subject can input the input vector composed of the character vector and the position encoding vector into the Bert layer to obtain the first vector that incorporates the context semantic relationship; then, the execution subject can input the above first vector into the Bi-LSTM layer to obtain the second vector that once again incorporates the context semantic relationship on the basis of the first vector; then, the execution subject can input the above second vector into the CRF layer to obtain the label of the character in the description text 401 and the probability corresponding to the character. Here, the label corresponding to "lace" indicated by icon 402 is BI, so it can be determined that "lace" is a selling point word. At this time, the next action can be executed (select "lace"). After executing this action, using the ctr=0.62 and cvr=0.12 corresponding to "lace", the corresponding reward value can be determined to be 0.47. Finally, we can use the strategy formula of the selling point word recognition model Determine the cumulative reward value and find an optimal strategy π so that the maximum long-term cumulative reward value can be obtained in any state s and any time t.

[0083] Further references Figure 5 As an implementation of the methods shown in the above figures, the present application provides an embodiment of a device for outputting information, and the device embodiment is similar to Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0084] like Figure 5As shown, the device 500 for outputting information in this embodiment includes: an acquisition unit 501 , an input unit 502 , a generation unit 503 and an output unit 504 . The acquisition unit 501 is configured to acquire the description text of the item, extract the character vector and position coding vector of the characters in the description text, wherein the position coding vector is used to represent the position of the character in the description text; the input unit 502 is configured to input the input vector into a pre-trained selling point word recognition model to obtain the label of the character in the description text and the probability of the character corresponding to the character, wherein the input vector is determined based on the character vector and the position coding vector of the character, the label includes a label used to represent that the character is a character in the selling point word, and the selling point word recognition model is based on the relevant indicators corresponding to the selling point words in the preset selling point word set, and is trained by using a reinforcement learning method, and the relevant indicators include at least one of the following: conversion rate and click-through rate; the generation unit 503 is configured to generate a selling point word set based on the labels of the characters in the description text; the output unit 504 is configured to determine the probability of the selling point word corresponding to the character in the selling point word, and select the selling point word from the selling point word set for output based on the probability of the selling point word corresponding to the character, and select the selling point word from the selling point word set for output based on the probability of the selling point word corresponding to the character.

[0085] In this embodiment, the specific processing of the acquisition unit 501, the input unit 502, the generation unit 503 and the output unit 504 of the device 500 for outputting information can be referred to. Figure 2 This corresponds to step 201, step 202, step 203 and step 204 in the embodiment.

[0086] In some optional implementations of the present embodiment, the selling point word recognition model generally includes a bidirectional encoding layer, a bidirectional long short-term memory layer, and a conditional random field layer. Among them, the bidirectional long short-term memory layer is generally composed of a forward LSTM and a backward LSTM, both of which are often used to model context information in natural language processing tasks. Conditional random field is a discriminant probability model, which is often used to annotate or analyze sequence data, such as natural language text or biological sequences. The above-mentioned input unit 502 can input the input vector into the bidirectional encoding layer of the above-mentioned selling point word recognition model to obtain a first vector. Here, the above-mentioned first vector is generally a character vector that incorporates contextual semantic relations. The above-mentioned input vector is generally determined based on the character vector and position encoding vector of the character. As an example, the character vector and the position encoding vector of the character can be spliced, and the splicing result is used as the input vector. Afterwards, the above-mentioned input unit 502 can input the above-mentioned first vector into the bidirectional long short-term memory layer of the above-mentioned selling point word recognition model to obtain a second vector. Here, the above-mentioned second vector is generally a character vector that once again incorporates contextual semantic relations on the basis of the first vector. Finally, the input unit 502 may input the second vector into the conditional random field layer of the selling point word recognition model to obtain the labels of the characters in the description text and the probability of the characters corresponding to the characters. The labels of the characters may include labels used to characterize that the characters are characters in the selling point words, and may also include labels used to characterize that the characters are not characters in the selling point words. Here, the labels may be marked in the form of {B, I, O}, that is, the characters at the beginning of the selling point words are usually marked as B, the characters in the selling point words are usually marked as I, and the other characters are usually marked as O. The probability of a character corresponding to a character usually refers to the probability that the character is the determined label.

[0087] In some optional implementations of this embodiment, the strategy of the above selling point word recognition model may be determined by the following formula (1):

[0088]

[0089] Among them, γ represents the discount rate, γ∈[0,1], γ k Usually, the current reward is more important than the future reward. k represents the number of executions, t represents the current moment, and r represents the current time. t+k Indicates the reward value obtained by the selling point word recognition model when it is executed for the kth time, s t Represents the current state. Represents the cumulative reward value obtained after performing a set of actions at the current state. represents the expected value of the cumulative reward value, argmax(f(x)) is the variable point x (or set of x) corresponding to the maximum value of f(x), and π represents the selling point word selection path corresponding to the maximum value of the expected value of the cumulative reward value. The optimization goal of the agent (the above selling point word recognition model) is to find an optimal strategy π so that the maximum long-term cumulative reward value can be obtained in any state s and any time t.

[0090] As an example, if the current state S t =(skuid=1,sp=lace,ctr=0.15,cvr=0.03), after executing the next action (selecting sp=knitting pattern), ctr=0.62,cvr=0.12, and the corresponding reward value can be determined based on ctr and cvr. The agent can select actions with different selling point words and select the action path with the maximum reward as the optimal solution.

[0091] In some optional implementations of this embodiment, the above-mentioned reward value may be a weighted average of the conversion rate and click-through rate corresponding to the selling point words, and the weights corresponding to the conversion rate and click-through rate may be preset experience values. Here, the reward value may be determined by the following formula (2):

[0092] r=0.7×ctr+0.3×cvr (2)

[0093] Among them, r represents the reward value, ctr represents the click rate, and cvr represents the conversion rate.

[0094] As an example, if ctr=0.62, cvr=0.12, the corresponding reward value may be 0.7×0.62+0.3×0.12=0.47.

[0095] In some optional implementations of this embodiment, the device 500 for outputting information may further include a determination unit (not shown in the figure). For each selling point word in the above-mentioned selling point word set, the above-mentioned determination unit may determine whether the selling point word corresponds to a relevant index. If there is no relevant index for the selling point word, the above-mentioned determination unit may obtain the k nearest neighbor selling point words of the selling point word. Here, the k nearest neighbor selling point words of the selling point word may be the k selling point words whose word vectors are closest to the word vector of the selling point word. The above-mentioned determination unit may determine the weighted average of the relevant indexes corresponding to the above-mentioned k nearest neighbor selling point words as the relevant index corresponding to the selling point word. Among them, for each neighbor selling point word in the above-mentioned k nearest neighbor selling point words, the weight of the neighbor selling point word may be the similarity between the neighbor selling point word and the selling point word. Here, the similarity between the neighbor selling point word and the selling point word may be the distance between the word vector of the neighbor selling point word and the word vector of the selling point word. The distance between word vectors may be a cosine distance or a Euclidean distance.

[0096] Here, the determination unit may determine the conversion rate of the selling point word by the following formula (3):

[0097]

[0098] Among them, cvr represents the conversion rate, i∈[1,k], cvr i represents the conversion rate corresponding to the i-th selling point word, sp represents the selling point word, sp i Represents the i-th selling point word, sim(sp i ,sp) represents the similarity between the selling point word and the i-th selling point word, and k represents the number of neighboring selling point words.

[0099] Here, the determination unit may determine the click rate of the selling point word by the following formula (4):

[0100]

[0101] Among them, ctr represents the click rate, i∈[1,k], ctr i represents the click-through rate corresponding to the i-th selling point word, sp represents the selling point word, sp i Represents the i-th selling point word, sim(sp i ,sp) represents the similarity between the selling point word and the i-th selling point word, and k represents the number of neighboring selling point words.

[0102] It should be noted that the word vectors of the selling point words and the neighboring selling point words can be determined based on the Word2vec model.

[0103] Reference below Figure 6 , which shows an electronic device (eg, Figure 1 A schematic diagram of the structure of the server in (a) 600. Figure 6 The server shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0104] like Figure 6 As shown, the electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0105] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 6 The electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 6 Each block shown in the figure may represent one device, or may represent multiple devices as required.

[0106] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, the above functions defined in the method of the embodiment of the present disclosure are executed. It should be noted that the computer-readable medium described in the embodiment of the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection with one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In an embodiment of the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, an apparatus, or a device. In an embodiment of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in combination with an instruction execution system, an apparatus, or a device. The program code embodied on the computer readable medium may be transmitted using any appropriate medium, including but not limited to: wire, optical cable, RF (radio frequency), etc., or any suitable combination of the foregoing.

[0107] The computer-readable medium may be included in the electronic device; or it may exist independently without being installed in the electronic device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: obtains the description text of the item, extracts the character vector and position coding vector of the characters in the description text, wherein the position coding vector is used to represent the position of the character in the description text; inputs the input vector into a pre-trained selling point word recognition model to obtain the label of the character in the description text and the probability of the character corresponding to the character, wherein the input vector is determined based on the character vector and the position coding vector of the character, the label includes a label used to represent that the character is a character in the selling point word, and the selling point word recognition model is based on the relevant indicators corresponding to the selling point words in the preset selling point word set, and is trained by using a reinforcement learning method, and the relevant indicators include at least one of the following: conversion rate and click-through rate; based on the labels of the characters in the description text, a selling point word set is generated; based on the probability of the characters in the selling point word corresponding to the selling point word, the probability of the selling point word corresponding to the selling point word is determined; and based on the probability of the selling point word corresponding to the selling point word, a selling point word is selected from the selling point word set for output.

[0108] Computer program code for performing the operations of embodiments of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on a user's computer, partially on a user's computer, as a separate software package, partially on a user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0109] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0110] The units involved in the embodiments described in the present disclosure may be implemented by software or hardware. The units described may also be set in a processor, for example, it may be described as: a processor includes an acquisition unit, an input unit, a generation unit, and an output unit. The names of these units do not constitute a limitation on the units themselves in some cases. For example, the generation unit may also be described as "a unit for generating a set of selling point words based on labels describing characters in a text".

[0111] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) to form a technical solution.

Claims

1. A method for outputting information, include: Obtaining a description text of an item, and extracting a character vector and a position encoding vector of a character in the description text, wherein the position encoding vector is used to represent a position of a character in the description text; Inputting an input vector into a pre-trained selling point word recognition model to obtain labels of characters in the description text and corresponding probabilities of the characters, wherein the input vector is determined based on a character vector and a position encoding vector of the character, the label includes a label used to characterize that the character is a character in a selling point word, and the selling point word recognition model is trained using a reinforcement learning method based on relevant indicators corresponding to selling point words in a preset selling point word set, and the relevant indicators include at least one of the following: conversion rate and click-through rate; Generate a set of selling point words based on the labels of the characters in the description text; Determining the probability of the selling point word corresponding to characters in the selling point word based on the probability of the selling point word corresponding to characters, and selecting a selling point word from the selling point word set for output based on the probability of the selling point word corresponding to characters; The selling point word recognition model includes a bidirectional encoding layer, a bidirectional long short-term memory layer and a conditional random field layer; and The step of inputting the input vector into a pre-trained selling point word recognition model to obtain labels of characters in the description text and corresponding probabilities of the characters includes: Inputting the input vector into the bidirectional coding layer to obtain a first vector; Inputting the first vector into the bidirectional long short-term memory layer to obtain a second vector; The second vector is input into the conditional random field layer to obtain labels of characters in the description text and corresponding probabilities of the characters.

2. The method according to claim 1, in, The strategy of the selling point word recognition model is determined by the following formula: Among them, γ represents the discount rate, γ∈[0,1], k represents the number of executions, t represents the current time, and r t+k Indicates the reward value obtained by the selling point word recognition model when it is executed for the kth time, s t Represents the current state. Represents the cumulative reward value obtained after performing a set of actions at the current state. Characterizes the expected value of the cumulative reward value, π Represents the selling point word selection path corresponding to the maximum expected value of the cumulative reward value.

3. The method according to claim 2, in, The reward value is a weighted average of the conversion rate and the click rate corresponding to the selling point words.

4. The method according to any one of claims 1 to 3, in, The method further comprises: For each selling point word in the set of selling point words, determine whether the selling point word corresponds to a relevant indicator; if not, obtain the k nearest neighbor selling point words of the selling point word, and determine the weighted average of the relevant indicators corresponding to the k nearest neighbor selling point words as the relevant indicator corresponding to the selling point word, wherein, for each neighbor selling point word in the k nearest neighbor selling point words, the weight of the neighbor selling point word is the distance between the word vector of the neighbor selling point word and the word vector of the selling point word.

5. A device for outputting information, include: An acquisition unit is configured to acquire a description text of an item, and extract a character vector and a position encoding vector of a character in the description text, wherein the position encoding vector is used to represent a position of a character in the description text; The input unit is configured to input an input vector into a pre-trained selling point word recognition model to obtain a label of a character in the description text and a probability of the character corresponding to the character, wherein the input vector is determined based on a character vector and a position encoding vector of the character, the label includes a label used to characterize that the character is a character in a selling point word, and the selling point word recognition model is trained by using a reinforcement learning method based on relevant indicators corresponding to selling point words in a preset selling point word set, and the relevant indicators include at least one of the following: conversion rate and click-through rate; A generating unit, configured to generate a set of selling point words based on labels of characters in the description text; an output unit configured to determine the probability of the selling point word corresponding to a character in the selling point word based on the probability of the character corresponding to the selling point word, and select a selling point word from the selling point word set for output based on the probability of the selling point word corresponding to the character; The selling point word recognition model includes a bidirectional encoding layer, a bidirectional long short-term memory layer and a conditional random field layer; and The input unit is further configured to input the input vector into a pre-trained selling point word recognition model in the following manner to obtain the labels of the characters in the description text and the corresponding probabilities of the characters: Inputting the input vector into the bidirectional coding layer to obtain a first vector; Inputting the first vector into the bidirectional long short-term memory layer to obtain a second vector; The second vector is input into the conditional random field layer to obtain labels of characters in the description text and corresponding probabilities of the characters.

6. The device according to claim 5, in, The strategy of the selling point word recognition model is determined by the following formula: Among them, γ represents the discount rate, γ∈[0,1], k represents the number of executions, t represents the current time, and r t+k Indicates the reward value obtained by the selling point word recognition model when it is executed for the kth time, s t Represents the current state. Represents the cumulative reward value obtained after performing a set of actions at the current state. represents the expected value of the cumulative reward value, and π represents the selling point word selection path corresponding to the maximum value of the expected value of the cumulative reward value.

7. An electronic device, include: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 4.

8. A computer readable medium having a computer program stored thereon, in, When the program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Push information generation method and device

    CN110245257A

  • Weight model training method and related device

    CN110276010A