Ad copy material extraction method and device, equipment, medium and product
By constructing query statements and filtering detailed statements based on similarity and confidence, the problem of insufficient copywriting ability among e-commerce platform store users has been solved, realizing efficient and personalized advertising copywriting assistance and improving the promotion effect of advertising.
Patent Information
- Application Number
- CN202210626061.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-02
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2042-06-02
AI Technical Summary
Many shop owners on e-commerce platforms lack professional copywriting skills, resulting in poor quality advertising copy and ineffective promotion. Existing technology cannot effectively provide personalized copywriting assistance.
By obtaining the title text and category tags of the advertised products, a query statement is constructed, and detailed statements that match the copy phrases are recalled. Based on similarity and confidence, suitable detailed statements are selected as copy materials to form a list of copy materials.
It improves the efficiency and expressiveness of advertising copywriting, ensuring that copywriting materials can comprehensively and accurately express product characteristics, and achieve personalized advertising copywriting assistance.
Smart Images

Figure CN114971730B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of e-commerce information, and in particular to a copywriting material extraction method and a corresponding device, computer equipment, computer readable storage medium, and computer program product. BACKGROUND
[0002] An e-commerce platform is usually configured with an advertisement launching page for a store user to launch an advertisement corresponding to a listed product in the store to a advertisement system, so as to achieve the purpose of online traffic and promote the transaction volume of the product.
[0003] When launching an advertisement, a corresponding advertisement copy is required. A professional copy usually has a better promotion effect. However, in reality, a large number of store users in the e-commerce platform do not have the ability to write a professional copy or cannot afford the high cost of writing services. However, the copy written by the store user himself cannot achieve an effective promotion effect due to the lack of professionalism.
[0004] The traditional processing method is to generate a corresponding advertisement copy by using a preset copy template based on the product specified by the store user. This method solves the problem of automatic generation of advertisement copy, but cannot reflect the personalized content of the store user. The phenomenon often occurs that the advertisement copy generated by the system cannot meet the subjective expectations of the store user, and the store user cannot grasp the pros and cons of the copy content written by himself.
[0005] Therefore, how to provide an effective advertisement copy auxiliary creation method for the advertisement launch of a product still has room for exploration. SUMMARY
[0006] The present application aims to solve the above problems and provide a copywriting material extraction method and a corresponding device, computer equipment, computer readable storage medium, computer program product,
[0007] To achieve the various purposes of the present application, the following technical solutions are adopted:
[0008] In one aspect, a copywriting material extraction method is provided to achieve one of the purposes of the present application, comprising:
[0009] Obtaining a title text and a category label of an advertisement product, and constructing a query statement;
[0010] Recalling a detail sentence in a detail text of the advertisement product according to a copywriting phrase matched with the title text and / or the category label;
[0011] Determining the similarity and confidence between the query statement and each matched detail sentence;
[0012] According to the similarity and the confidence, part of the detail sentences are screened out as the copy materials of the advertisement copy of the advertisement commodity, and a copy material list is constituted.
[0013] On the other hand, a copy material extraction device is provided for one of the purposes of the present application, comprising a query construction module, a sentence recall module, a matching processing module, and a material generation module, wherein: the query construction module is used to acquire the title text and the category label of an advertisement commodity, and construct a query sentence; the sentence recall module is used to recall detail sentences in the detail text of the advertisement commodity according to the copy phrases matched with the title text and / or the category label; the matching processing module is used to determine the similarity and the confidence between the query sentence and each of the matched detail sentences; and the material generation module is used to screen out part of the detail sentences as the copy materials of the advertisement copy of the advertisement commodity according to the similarity and the confidence, and constitute a copy material list.
[0014] In another aspect, a computer device is provided for one of the purposes of the present application, comprising a central processing unit and a memory, and the central processing unit is used to call and run a computer program stored in the memory to execute the steps of the copy material extraction method described in the present application.
[0015] In another aspect, a computer readable storage medium is provided for another purpose of the present application, which stores a computer program realized according to the copy material extraction method in the form of computer readable instructions, and when the computer program is called and run by a computer, the steps included in the method are executed.
[0016] In another aspect, a computer program product is provided for another purpose of the present application, comprising a computer program / instruction, and when the computer program / instruction is executed by a processor, the steps of the copy material extraction method described in any one of the embodiments of the present application are realized.
[0017] Compared with the prior art, the present application has many advantages, at least including the following aspects:
[0018] Firstly, in one aspect of the present application, the query sentence is constructed by using the commodity title and the category label of the advertisement commodity to be published to the advertisement system, and in another aspect, the copy phrases are recalled according to the commodity title and / or the category label, and the detail sentences in the detail text of the advertisement commodity are acquired according to the recalled copy phrases, and then, the information contribution value of each detail sentence is determined comprehensively according to the similarity and the confidence between the query sentence and each acquired detail sentence, and a part of the detail sentences are screened out as the copy materials, so as to ensure that the recalled copy materials are the text contents capable of expressing the commodity characteristics of the advertisement commodity, and the writing efficiency and the expression ability of the advertisement copy of the advertisement commodity can be improved, and the advertisement copy assisted creation is realized.
[0019] Secondly, in the process of preparing the copy material in the application, the similarity and the confidence between the query statement and the detail statement are determined synchronously according to the semantic relationship between the query statement and the detail statement, wherein the similarity indicates the correlation degree between the product title and the category label in the advertising product and the detail statement, which can represent the close degree of the description of the product features of the advertising product by the detail statement, and the confidence is mainly based on the detail statement and can be used to indicate whether the statement is suitable as a promotion copy. The copy material screened out with reference to the similarity and the confidence can effectively represent the information contribution value of each detail statement to the advertising copy of the advertising product, facilitating the measurement of the advantages and disadvantages of each detail statement, so that the high-quality statements in the product detail text of the advertising product can be effectively selected.
[0020] In addition, the detail statements in the detail text of the advertising product are matched according to the copy phrases, and the copy phrases are matched according to the product title and the category label. Therefore, the copy phrases play a role in expanding the semantics of the product title and the category label of the advertising product, so that the more comprehensive copy phrases matched can ensure that more comprehensive detail statements are matched from the detail text in the subsequent process, realizing data checking, ensuring that effective detail statements are not missed, and then screening according to the similarity and the confidence corresponding to each detail statement, so that data checking can be realized. Therefore, the determined copy material is comprehensive and accurate. BRIEF DESCRIPTION OF DRAWINGS
[0021] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description, taken in conjunction with the following drawings, in which:
[0022] Figure 1 The flowchart of a typical embodiment of the copy material extraction method of the present application.
[0023] Figure 2 The flowchart of the detail statement recall process in the embodiment of the present application.
[0024] Figure 3 The flowchart of constructing the phrase library in the embodiment of the present application.
[0025] Figure 4 The flowchart of determining the similarity and the confidence of the detail statement in the embodiment of the present application.
[0026] Figure 5 The network architecture diagram of the exemplary text matching classification model of the present application.
[0027] Figure 6 The flowchart of the training process of the text matching classification model in the embodiment of the present application.
[0028] Figure 7 A flowchart of a process of selecting a script material in an embodiment of the present application.
[0029] Figure 8 A flowchart of a process of publishing an advertisement in an embodiment of the present application.
[0030] Figure 9 A principle block diagram of a script material extraction device of the present application;
[0031] Figure 10 A structure diagram of a computer device used in the present application. DETAILED DESCRIPTION
[0032] Embodiments of the present application are described in detail below with reference to examples illustrated in the accompanying drawings, in which the same or similar components have the same or similar designations and functions throughout various figures and / or texts. The embodiments described below with reference to the accompanying drawings are exemplary and are for the purpose of explanation only, and should not be construed as limiting the present application.
[0033] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a," "an," and "the" as used herein are intended to include plural forms as well. It should be further understood that the terms "comprises" and / or "comprising," when used in the specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It will be understood that when an element is referred to as being "connected" or "coupled" to another element, it can be directly connected or coupled to the other element or intervening elements can also be present. In addition, the term "connected" or "coupled" as used herein can include the case of wireless connection or wireless coupling. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0034] Those skilled in the art can understand that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. It should also be understood that the terms, such as those defined in a generally used dictionary, should be interpreted as having a meaning that is consistent with the meaning in the context of the relevant art and should not be interpreted in an idealized or overly formal sense unless specifically so defined herein.
[0035] Those skilled in the art will understand that, as used herein, the terms "client," "terminal," and "terminal device" include both devices that are solely wireless signal receivers and devices that have both receiving and transmitting hardware that can communicate bi-directionally over a bi-directional communication link. Such devices can include cellular or other communication devices with single-line or multiple-line displays, or no display, Personal Communications Service (PCS) devices that can combine a voice and / or data processor, a PDA that can include a radio frequency receiver and a pager, Internet and / or Intranet access, a Web browser, a calendar, and / or a GPS receiver, a conventional laptop and / or palmtop computer and / or other devices that have a radio frequency receiver. As used herein, the terms "client," "terminal," and "terminal device" can be portable, transportable, installed in a vehicle (aeronautical, maritime, and / or land), or adapted and / or configured for local and / or distributed operation on Earth and / or any other location in space. As used herein, the terms "client," "terminal," and "terminal device" can also be a communication terminal, an Internet terminal, a music / video playing terminal, such as a PDA, a Mobile Internet Device (MID), and / or a mobile phone with music / video playing function, a smart television, a set-top box, and / or the like.
[0036] As used herein, the terms "server," "client," "service node," and the like refer to hardware that has the equivalent capability of a personal computer, i.e., an electronic device having a central processing unit (including an arithmetic unit and a controller), a memory, an input device, and an output device, and the like necessary components disclosed by the Von Neumann principle. A computer program is stored in the memory, the central processing unit loads the program stored in the external memory into the memory and runs it, executes the instructions in the program, and interacts with the input and output devices, thereby completing a specific function.
[0037] It should be noted that the concept of "server" in the present application can also be extended to the case of a server cluster. According to the principle of network deployment understood by those skilled in the art, the servers should be logically divided, and in physical space, these servers can be independent of each other but can be called through an interface, or can be integrated into a physical computer or a computer cluster. Those skilled in the art should understand this variation and should not be restricted by the implementation of the network deployment of the present application.
[0038] One or more technical features of the present application, unless explicitly specified, can be deployed on a server for implementation and accessed by a client remotely calling an online service interface provided by the server, or can be directly deployed and run on a client for implementation.
[0039] The neural network model referred to or possibly referred to in the present application, unless explicitly specified, can be deployed on a remote server and remotely called by a client, or can be deployed on a client with sufficient device capability for direct calling. In some embodiments, when it runs on a client, its corresponding intelligence can be obtained through transfer learning to reduce the requirement for client hardware running resources and avoid excessive occupation of client hardware running resources.
[0040] The various data involved in the present application, unless explicitly specified, can be stored remotely on a server or stored locally on a terminal device, as long as it is suitable for being called by the technical solutions of the present application.
[0041] Those skilled in the art should know that the various methods of the present application, although based on the same concept and described to present commonality among them, are independently executable unless otherwise specified. Similarly, for each embodiment disclosed in the present application, it is based on the same inventive concept, so the same concept is understood to be equivalent, and although the concept is expressed differently, it is only a suitable transformation for convenience.
[0042] Unless it is explicitly stated that the embodiments disclosed in the present application are mutually exclusive, the technical features involved in each embodiment can be combined flexibly to construct new embodiments, as long as such combination does not deviate from the spirit of the present application and can meet the needs of the prior art or solve some deficiencies in the prior art. For this variation, those skilled in the art should know.
[0043] The text material extraction method of the present application can be programmed as a computer program product, deployed in a client or server for running and implementation, for example, in the exemplary application scenario of the present application, it can be deployed and implemented in the server of an e-commerce platform, thereby the interface opened after running of the computer program product can be accessed, the process of the computer program product is interacted with the human-computer interface to execute the method.
[0044] Please refer to Figure 1 In a typical embodiment of the text material extraction method of the present application, the following steps are included:
[0045] Step S1100, obtaining the title text and category label of the advertising goods, and constructing a query statement;
[0046] The e-commerce platform is usually equipped with an advertising system, which opens a corresponding advertising publishing interface to each online store merchant user, obtains the advertising copy and advertising configuration information corresponding to the advertising launched by the merchant user of any store to the advertising system through the advertising publishing interface, and submits it to the advertising publishing channel of the advertising system or the advertising publishing channel of a third party to display to the public.
[0047] In the present application, the advertising copy mainly includes advertising text, which is expressed by natural language and can be any language. Each store can be deployed in an independent site of the e-commerce platform. Each store can have a large number of goods on the shelf, and users can access the transaction page of any goods in the store to place orders and pay, etc., thereby realizing e-commerce transactions. Each store can launch advertising corresponding to any goods in its store to the advertising system, such goods can be called advertising goods, and the advertising copy corresponding to the advertising goods is provided in the advertising publishing process and submitted to the advertising system through the advertising publishing interface for publishing.
[0048] The goods information database of the online store stores the goods information of each goods on the shelf, including but not limited to the category label to which the goods belongs, the goods title which plays a main prompting role, the detail text which plays a comprehensive introduction role of the goods information, etc. Among them, the goods title is generally a combination of relatively concise words, usually composed of multiple keywords corresponding to the characteristics of the goods, and the goods detail text is relatively detailed in length, describing the characteristics of the goods from different aspects and details.
[0049] The category label is an identifier corresponding to the specific category of the commodity in the preset category system of the e-commerce platform. Specifically, the store of the e-commerce platform usually has a category system of commodities, which is used to classify and summarize the massive commodities in the store. The category system can be a multi-level classification system, that is, it contains multiple classification levels, and each classification level contains multiple specific categories. The construction of the category system can be provided by the e-commerce platform as a template, which is revised and determined by the merchant user.
[0050] In the advertisement script, the corresponding language can be used to describe the product characteristics and other advertisement information of the advertised product. The product characteristics can include the name, brand, features, functions, parameters, attributes, and any other information that helps the audience understand the product. In one embodiment, the text content used to describe the product characteristics can be a separate sentence in the product details text, that is, a detail sentence, or a replacement sentence further optimized and edited based on the detail sentence without changing its main meaning.
[0051] When the merchant user needs to make an advertisement script for the target product of the online store to achieve the purpose of advertisement publishing, the advertisement publishing interface provided by the e-commerce platform can be called to specify the target product as an advertised product. After the advertisement system receives the advertised product specified by the user on the server side, it can call the product title, category label, and detail text of the advertised product from the product information library. Further, the product title and category label are spliced to form a query statement, which is used to recommend advertisement script materials for the advertised product in the future. In one embodiment, the product title can be pre-processed before the query statement is constructed, such as removing stop words and invalid characters. As for the splicing order of the product title and the category label, it can be set as needed, as long as the uniform splicing rule is maintained.
[0052] Step S1200, according to the title text and / or the category label matched with the script phrase, recalling the detail sentence in the detail text of the advertised product matched with the script phrase;
[0053] The application provides a corresponding phrase library for each category in the category system. The phrase library corresponding to the category can be determined through the category label of the advertised product, and each phrase library collects the script phrases from the advertisement scripts of historical advertisements. In one embodiment, the script phrase can be composed of two or more word elements, which can be keywords with the same or different parts of speech in the historical advertisement scripts of the products belonging to the same category as the advertised product. The part of speech can be any one of a noun plus a noun, a verb, or an adjective. It is not difficult to understand that the script phrase is a partial phrase selected from the historical advertisement scripts, which has a certain information reference function.
[0054] Further, according to any one of the title text, category label or combination thereof, a plurality of text phrases can be matched from the phrase library, which can generally be obtained by semantic matching. In an embodiment, the phrase library is implemented by multi-channel recall with the title text and category label respectively, and the feature vectors corresponding to the title text and category label are obtained by using a text feature extraction model respectively, and then the feature vectors are respectively matched with each text phrase in the phrase library to determine the text phrases with a matching degree (similarity) reaching a preset threshold, thereby realizing the recall of the text phrases.
[0055] The essence of realizing the recall of the text phrases is to expand the semantic range of the advertisement commodity based on the commodity title and category label of the advertisement commodity, so the text phrases can be further used to match the detail sentences in the detail text of the advertisement commodity, so as to realize the full search of the detail sentences in the detail text of the advertisement commodity by using the expanded semantic range.
[0056] In an embodiment, in order to match the detail sentences in the detail text of the advertisement commodity according to the text phrases, the detail text can be divided into a plurality of independent natural sentences, i.e. detail sentences, by using natural punctuation marks as a separation mark, to form a sentence list for matching. Based on the sentence list, the semantic matching between each text phrase and each detail sentence in the sentence list can be realized to recall the matching detail sentences.
[0057] It is not difficult to understand that when each text phrase matched from the phrase library is used to match the corresponding detail sentence from the detail text, the semantic matching detail sentence is actually obtained according to the commodity title and / or category label, and although the two are mediated by the text phrase, they still have semantic relevance.
[0058] Step S1300, determining the similarity and confidence between the query sentence and each matched detail sentence;
[0059] The close degree of the association between the query sentence composed of the commodity title and category label and each detail sentence recalled from the detail text can be represented by the semantic similarity between the two, so that the higher the similarity, the more the corresponding detail sentence echoes the commodity title, which can better represent the commodity characteristics of the corresponding advertisement commodity; the lower the similarity, the lower the echo degree of the corresponding detail sentence and the commodity title, thereby relatively failing to effectively describe the commodity characteristics of the advertisement commodity.
[0060] In an embodiment, a plurality of categories representing different close degree levels are allowed to be set for the determination of the similarity, which can be mapped to different categories according to the similarity, and subsequently, the detail sentences with lower similarity can be quickly filtered by category screening.
[0061] The determination of the similarity can be implemented by a neural network model, and the neural network model is preferably a recurrent neural network model (RNN) such as LSTM (long short-term memory recurrent neural network), BiLSTM (bidirectional long short-term memory recurrent neural network), Transformer, Bert, SimCSE, etc. According to the principle disclosed in the present application, the neural network model can be trained to a convergent state in advance by using a sufficient amount of corresponding training samples, so that it learns the ability to determine the similarity according to the given query statement, the detail statement or their combination.
[0062] To this end, in one embodiment, two isomorphic basic neural network models can be used to build a double-tower model, and after the two basic neural network models extract feature vectors from the query statement and the detail statement respectively and splice the comprehensive feature vectors, further classification mapping is performed to obtain the classification probability mapped to the preset category as the similarity; in another embodiment, a single basic neural network model can be used to extract features from the combined text of the query statement and the detail statement to obtain a comprehensive feature vector, and then perform classification mapping to obtain the classification probability mapped to the preset category as the similarity.
[0063] The confidence between the query statement and each detail statement recalled is mainly used to represent whether the detail statement is suitable for promotion. The higher the confidence, the higher the information contribution value of the corresponding detail statement to the advertising copy of the advertising product. The lower the confidence, the lower the information contribution value of the detail statement to the advertising copy of the advertising product.
[0064] The determination of the confidence can also be implemented by a neural network model, and the neural network model is preferably a recurrent neural network model (RNN) such as LSTM (long short-term memory recurrent neural network), BiLSTM (bidirectional long short-term memory recurrent neural network), Transformer, Bert, SimCSE, etc. According to the principle disclosed in the present application, the neural network model can be trained to a convergent state in advance by using a sufficient amount of corresponding training samples, so that it learns the ability to determine the confidence according to the given query statement, the detail statement or their combination.
[0065] To this end, in one embodiment, two isomorphic base neural network models can be employed to build a double tower model, and after extracting feature vectors from the query statement and the detail statement respectively by the two base neural network models and splicing the comprehensive feature vectors to obtain the classification probability mapped to the preset category as the confidence after further classification mapping, in another embodiment, a single base neural network model can be used to extract features from the combined text of the query statement and the detail statement to obtain a comprehensive feature vector, and then perform classification mapping to obtain the classification probability mapped to the preset category as the confidence.
[0066] According to the above implementation principle of similarity and confidence, it can be understood that the determination of similarity and confidence can be implemented synchronously, so in one embodiment, it can be determined in parallel, that is, the similarity and confidence are determined synchronously for each detail statement, thereby improving the determination efficiency of similarity and confidence. Moreover, in fact, the determination of similarity and confidence can be determined based on the same model architecture, so in one embodiment, the neural network model architecture for calculating similarity and confidence adopts the same neural network model architecture, but after obtaining the comprehensive feature vector of the query statement and the detail statement, it is divided into two branches, and classification mapping based on similarity and confidence is performed respectively to obtain the corresponding similarity and confidence.
[0067] It is not difficult to understand that such a model architecture can be prepared by joint training to a convergent state, so that it can learn the ability to determine the corresponding similarity and confidence of the query statement and the detail statement simultaneously. In the process of joint training, the loss value of each branch is calculated adaptively using the corresponding supervision label of similarity and confidence, and then the sum of the loss values of the two branches is used to implement gradient update of the entire model architecture. As can be seen, for the joint input of the query statement and the detail statement, the similarity and confidence between them can be determined simultaneously, thereby determining the semantic close degree of the detail statement and the product title, category label in the query statement, and the extendable degree of the detail statement for advertising.
[0068] In order to enable the similarity to represent the semantic association close degree between the query statement and the detail statement, pre-labeled training samples can be used, and the corresponding supervision label can be set according to the actual association between the query statement and the detail statement in the training sample.
[0069] Similarly, in order to enable the confidence to represent the extendable degree of the detail statement, the training sample can also be manually labeled in advance, and the corresponding supervision label can be given according to whether the corresponding advertising copy of the training sample contains information that can meet the consumer's concern about some benefits.
[0070] Therefore, for each detail sentence of each recall, the similarity and the confidence corresponding to each detail sentence can be determined according to the above principles by combining the query sentence.
[0071] In step S1400, some detail sentences are filtered according to the similarity and the confidence as the copy materials of the advertisement copy of the advertisement commodity, to form a copy material list.
[0072] After each detail sentence of each recall obtains the corresponding similarity and confidence, the similarity and the confidence can be further used to sort the detail sentences of the recall. There are various ways to sort. For example:
[0073] In one embodiment, for each detail sentence, the weighted result of the similarity and the confidence is obtained, and the weighted result is used to sort the detail sentences.
[0074] In another embodiment, the similarity is used as the primary index, and the confidence is used as the secondary index, to sort the detail sentences by multiple indexes.
[0075] In any case, the sorting of the detail sentences of the recall can be realized with the help of the similarity and the confidence. After sorting, according to the expected preset number, the corresponding part of the detail sentences with high priority can be selected. It is not difficult to understand that the detail sentences with high priority are generally the detail sentences with relatively high similarity and confidence. These detail sentences are used as the copy materials required by the advertisement copy, to form a copy material list, which can be pushed to the terminal device of the user for display, so that the user can refer to it to create the corresponding advertisement copy.
[0076] According to the above embodiments, it can be seen that the present application has many advantages, at least including the following aspects:
[0077] Firstly, the present application uses the product title and the category label of the advertisement commodity to be published to the advertisement system to construct a query sentence, and recalls the copy phrases according to the product title and / or the category label, and obtains the detail sentences from the detail text of the advertisement commodity according to the recalled copy phrases. Then, the information contribution value corresponding to each detail sentence is determined according to the similarity and the confidence between the query sentence and each obtained detail sentence, and a part is filtered as the copy materials, so as to ensure that the recalled copy materials are the text content that can express the product features of the advertisement commodity, which can improve the writing efficiency and expression ability of the advertisement copy of the advertisement commodity, and realize the auxiliary creation of the advertisement copy.
[0078] Secondly, in the process of preparing the copy material in the application, the similarity and the confidence between the query statement and the detail statement are determined synchronously according to the semantic relationship between the query statement and the detail statement, wherein the similarity indicates the correlation degree between the product title and the category label in the advertising product and the detail statement, which can represent the close degree of the description of the product features of the advertising product by the detail statement, and the confidence is mainly based on the detail statement and can be used to indicate whether the statement is suitable as a promotion copy. The copy material screened out with reference to the similarity and the confidence can effectively represent the information contribution value of each detail statement to the advertising copy of the advertising product, facilitating the measurement of the advantages and disadvantages of each detail statement, so as to effectively select the high-quality statements in the product detail text of the advertising product.
[0079] In addition, the detail statements in the detail text of the advertising product are matched according to the copy phrases, and the copy phrases are matched according to the product title and the category label. Therefore, the copy phrases play a role in expanding the semantics of the product title and the category label of the advertising product. The content-rich copy phrases matched in this way can ensure that more comprehensive detail statements are matched from the detail text in the subsequent process, achieving data completeness and ensuring that effective detail statements are not missed. Subsequent screening according to the similarity and the confidence corresponding to each detail statement can achieve data precision. Therefore, the determined copy material is both comprehensive and accurate.
[0080] According to the extended embodiments of any of the above embodiments, referring to Figure 2 , the step S1200 of recalling the detail statements in the detail text of the advertising product matched with the copy phrases according to the copy phrases matched with the title text and / or the category label comprises:
[0081] The step S1210 of segmenting the detail text of the advertising product to obtain a sentence list composed of the detail statements in the detail text;
[0082] The detail text of the advertising product can contain interference information such as HTML tags, expression characters, punctuation marks, etc. The interference information or other interference information can be removed by data cleaning, and then the tokenizer function for realizing segmentation provided by NLTK (Natural Language Tool Kit) is used to segment the data cleaned detail text, so as to obtain a sentence list containing multiple independent statements extracted from the detail text. These independent statements are the detail statements.
[0083] Step S1220, according to the title text and / or the category label, a plurality of copy phrases are matched from the phrase library corresponding to the category label, and a phrase list is constituted, wherein the copy phrases include a plurality of word units with independent word types;
[0084] In this embodiment, recall operations corresponding to two or more recall channels can be implemented from the phrase library corresponding to the category label of the advertised product, and these recall operations include recall according to the title text alone, recall according to the category label alone, and recall according to the title text and the category label respectively.
[0085] Before implementing the multi-channel recall, the copy phrases in the phrase library, the product title of the advertised product, and the category label and other to-be-processed texts can be subjected to word embedding and converted into high-dimensional feature vectors by means of various pre-trained neural network models provided by the open source framework Sentence Transformers, including but not limited to BERT, RoBERTa, XLM-RoBERTa, MPNet, etc.
[0086] On the basis that the copy phrases, the product title of the advertised product, and the category label all have their feature vectors, a preset data distance algorithm can be used to determine the similarity between each other by calculating the data distance between the product title, the category label, and each copy phrase. The data distance algorithm can be any one or any multiple of the cosine similarity algorithm, the Euclidean distance algorithm, the Jaccard coefficient algorithm, and the Pearson coefficient algorithm. After determining the similarity of each copy phrase in the phrase library in each recall pass, a part of the copy phrases with higher similarity can be selected according to a preset threshold or a preset number to constitute a phrase list. Each copy phrase in the phrase list is a copy phrase recalled according to the product title and the category label. As mentioned above, the copy phrase includes a plurality of word units with independent word types.
[0087] Step S1230, the similarity of each copy phrase in the phrase list to each detail sentence in the sentence list is calculated, and a detail sentence that constitutes semantic matching with each copy phrase is selected according to the similarity.
[0088] After the semantic expansion according to the product title and the category label of the advertised product and the obtaining of the phrase list according to the semantic range after the semantic expansion, each copy phrase in the phrase list can be used to recall a part of the detail sentences from the sentence list corresponding to the detail text of the advertised product, so as to realize data check for the detail sentences in the detail text.
[0089] Similarly, each detail sentence in the sentence list can also generate a high-dimensional feature vector by the same neural network model as the previous step, and then, similarly to the operation of recalling the advertising phrase in the previous step, the data distance between each advertising phrase in the phrase list and each detail sentence in the sentence list is calculated to determine the corresponding similarity, and then the detail sentences with high similarity to each advertising phrase are screened out. These recalled detail sentences are the detail sentences that are semantically matched with the advertising phrases in the phrase list.
[0090] According to the above embodiments, in the process of recalling the detail sentences corresponding to the advertising goods from the detail text of the advertising goods, part of the advertising phrases are recalled by means of the phrase library corresponding to the category label of the advertising goods. These advertising phrases are semantically matched with the product title and category label of the advertising goods, so that the semantic expansion of the product title and category label of the advertising goods is realized through these advertising phrases. Then, the detail sentences are further recalled from the detail text of the advertising goods according to these advertising phrases, so as to avoid missing key information and realize data check. In this process, semantic matching can be based on feature vectors, and the operation efficiency is relatively high.
[0091] According to the extended embodiments of any of the above embodiments, please refer to Figure 3 , before the step S1220, according to the advertising phrases matched with the title text and / or category label, includes:
[0092] Step S2100, extracting a plurality of advertising phrases from the advertising texts of the advertising goods corresponding to the category label in the advertising system to form a candidate phrase, the advertising phrases are extracted according to a plurality of preset phrase structures, and the phrase structure includes a plurality of ordered part-of-speech tags, at least one of which contains a part-of-speech tag representing a noun, which is postfixed relative to other part-of-speech tags;
[0093] In this application, the phrase library corresponding to each category of the category system of the e-commerce platform, the advertising phrases in which can be extracted and prepared from the advertising texts of the same category of goods in the historical advertising system associated with the e-commerce platform. For this purpose, for each online store, the advertising texts required by the phrase library corresponding to each category label of the store are the advertising texts used by the goods corresponding to the category label. The advertising phrases extracted from these advertising texts according to the preset rules can be used to construct the phrase library of the corresponding category.
[0094] In one embodiment, part-of-speech structure information is provided in advance, which is used to define the word structure rules of the advertising phrases to be extracted from the advertising texts, so it can be represented by including a plurality of phrase structures. For example, the phrase structure is represented in the following form:
[0095] Noun & Noun
[0096] Adjective & Noun
[0097] Verb & Noun
[0098] As can be seen, each phrase structure includes a plurality of word class labels arranged in order, for indicating that words of the same (Noun & Noun) or different word classes (Adjective & Noun, Verb & Noun) are combined into a text phrase. In view of the habit in natural language that a commodity noun is postposed relative to adjectives, verbs, etc., the word class label corresponding to the noun can also be placed at the end position in the order in the phrase structure.
[0099] For each advertising text, in order to obtain a text phrase therefrom, a preset word segmentation manner can be applied first, such as using an N-Gram algorithm, a Jieba word segmenter, etc. to perform word segmentation, and meanwhile, a preset word class extractor or other preset neural network model for realizing word class labeling is used to perform word class labeling on each word segmentation, to obtain the word class corresponding to each word segmentation.
[0100] In an embodiment, the neural network model for realizing word class labeling can be implemented by using a text feature extractor such as LSTM (Long Short-Term Memory) or Bert (Bidirectional Encoder Representation from Transformers) in combination with a conditional random field CRF.
[0101] In an embodiment, when performing word segmentation on the advertising text, a sliding window of a smaller size can be used to take words first, such as using two single-character lengths, where the single character is a single Chinese character for Chinese, or a single word for a phonetic language such as English. By taking words with a smaller sliding window to obtain the word segmentation of the advertising text, word class labeling is performed thereon, and the word segmentation with the word class labeled can be used as a word unit.
[0102] In an embodiment, the N-Gram algorithm can be used to perform word segmentation on the advertising text multiple times by increasing the sliding window size, such as setting the sliding window to four-character words or five-character words, to obtain a plurality of candidate phrases in the advertising text. As can be understood, these candidate phrases can include word segmentations (two-character words, three-character words, etc.) obtained with a smaller sliding window, and the word classes of these word segmentations have been determined.
[0103] In another embodiment, according to the phrase structure and the specification of the combination relationship of different word classes in the phrase structure, each word segmentation can be combined as a word unit in sequence to obtain a plurality of candidate phrases.
[0104] For each advertising copy, after obtaining its corresponding multiple candidate phrases, each candidate phrase is matched one by one according to the phrase construction in the word class structure information. When the word combination relationship of a candidate phrase matches a phrase construction, the candidate phrase is determined as a copy phrase, otherwise the candidate phrase is discarded. After matching with the candidate phrases, the corresponding copy phrases can be matched and determined from the multiple candidate phrases of an advertising copy, and can be stored in the phrase library of the corresponding category. In an embodiment, before matching the phrase construction, each candidate phrase can be further de-duplicated. The word stem of each candidate phrase is extracted, and then the candidate phrases with the same word stem are de-duplicated, and only one is retained.
[0105] Step S2200, determine the information contribution score of each candidate phrase according to the category, store, and advertisement where the candidate phrase is located;
[0106] In order to clarify the information contribution value of the copy phrase in the advertising information, the information contribution score of the copy phrase can be determined. The information contribution score is essentially a recommendation degree, which can be determined by referring to the advertising copy of the already launched advertisement obtained from the advertising system. The information reference value of each copy phrase is quantified through the information contribution score.
[0107] When determining the information contribution score, the information contribution score of each copy phrase is determined according to the association of the category, store, and advertisement corresponding to the copy phrase.
[0108] For the category dimension, since the frequency of use of each copy phrase in the advertising copy of the goods in each category is different, for a category, the frequency of use of each copy phrase under the category in the advertising copy of the category is also different, that is, in the phrase library corresponding to a category, the frequency of use of each copy phrase is also different, which means that the popularity of each copy phrase is different. According to this principle, for each phrase library, the category dimension score corresponding to each copy phrase can be quantitatively determined, which is used to represent the information contribution value of the copy phrase in the advertising copy of the corresponding category.
[0109] For the store dimension, each store has different preferences for using the same copy phrase in the advertising copy of the same category of goods, so even if each copy phrase in the phrase library corresponding to the same category has a higher frequency in one store and a lower frequency in other stores to which the advertising copy of the same category belongs, the freshness of the copy phrase obtained in the former store is obviously higher than that in the latter store. Therefore, even if the same copy phrase and the same category, the information contribution value corresponding to different stores is different. According to this principle, for each copy phrase in the phrase library of each category, the store dimension score of each store can be quantitatively determined to represent the information contribution value of the copy phrase in the store publishing the advertising copy of the same category to a specific store.
[0110] For the advertising dimension, since each advertisement will generate corresponding effectiveness data, each advertisement uses a corresponding advertising copy, and each advertising copy contains one or more copy phrases stored in the phrase library corresponding to the category of the goods corresponding to the advertising copy. Therefore, each copy phrase in each phrase library can obtain corresponding effectiveness data according to the advertising copy to which it belongs. By statistically analyzing these effectiveness data, the advertising dimension score corresponding to each copy phrase can be obtained to represent the information contribution value of the copy phrase in the effectiveness data of the advertising copy of the same category. The advertising effectiveness data includes but is not limited to: click-through rate (CTR), collection rate, add-to-cart rate, conversion rate (CVR), return on advertising spend (ROAS), etc.
[0111] After obtaining the category dimension score, the store dimension score, and the advertising dimension score of the copy phrase, the information contribution score of each copy phrase in the phrase library of each category can be determined for each store. For example, for each store, when calculating the information contribution score of each copy phrase in each phrase library, the category dimension score corresponding to the copy phrase, the store dimension score of the copy phrase relative to the current store, and the advertising dimension score corresponding to the copy phrase are summarized. The summary method can use any method such as summation, mean value, weighted sum, etc. Thus, the information contribution score of the copy phrase in the current store can be obtained.
[0112] It can be seen that for each copy phrase, the information contribution score obtained is different as long as the given store and category are different, that is, the embodiment adopts a unified processing process to complete the information contribution score obtained by each copy phrase under a specific category and a specific store, thereby realizing the personalized customization of the copy phrases of each store. In addition to being affected by the phrase library divided by category, it is mainly affected by the store dimension score, so that each copy phrase is associated with the store's preference for the copy phrase to determine its final information contribution score.
[0113] As can be seen, the information contribution score of the copy phrase gives the information contribution value of the category, store, advertisement and other aspects, and realizes personalized corresponding scoring with the store, which has the effect of efficiently representing the actual information value.
[0114] Step S2300, filtering part of the candidate phrases according to the information contribution score, and retaining the copy phrases stored as the phrase library.
[0115] Each copy phrase in the phrase library has its corresponding information contribution score, but the score is high or low, and the candidate phrases determined in the foregoing can be selected according to a preset threshold or a preset number, and a plurality of candidate phrases with information contribution scores higher than the preset threshold or according to the top preset data of the information contribution score are selected as the final copy phrases. The final copy phrases are used to construct the phrase library.
[0116] The above embodiment extracts copy phrases composed of two or more word units from the advertising copy of the historical advertisements of the advertising system, constructs a phrase library, and then combines the category dimension, the store dimension and the advertisement dimension corresponding to the advertising effect to comprehensively score each copy phrase, obtains the information contribution score of each copy phrase under the constraint of the store and the category, and finally filters the copy phrases according to the information contribution score, realizes the optimization of the copy phrases, and obtains the final phrase library. The copy phrases in the phrase library have the function of representing high-value word choice and sentence construction in the historical advertising copy, and can provide semantic reference for subsequent determination of detail sentences.
[0117] In an embodiment, the category can be used as a dimension, and the category dimension score of each copy phrase can be determined according to the occurrence proportion of the copy phrase in the advertising copy of the same category;
[0118] First, the word frequency of each copy phrase in the advertising copy of the same category of goods and the number of advertising copy of the same category of goods are counted, and the ratio of the word frequency to the number of advertising copy is determined as the occurrence proportion of the corresponding copy phrase;
[0119] First, take each category as an independent unit, count the occurrence frequency of each copy phrase w in the phrase library corresponding to each category j in all the advertising copy of the advertising that has been launched, that is, the word frequency freqency w_j .
[0120] Then count the number of advertising copy count of all the advertising copy of the advertising that has been launched corresponding to the category j , so that the occurrence proportion Ratio of each copy phrase in all the advertising copy of the advertising that has been launched can be obtained w_j , that is:
[0121] Ratio w_j = freqency w_j / count j
[0122] Then, the occurrence proportion of all copy phrases corresponding to each category is normalized by category to obtain the category dimension score of each copy phrase under the corresponding category.
[0123] The occurrence proportion of all copy phrases in the phrase library of each category can be normalized to realize numerical specification and adjust the statistical dimension of each occurrence proportion to the numerical space of [0, 1]. In an embodiment, the softmax function is applied for normalization to convert the occurrence proportion of each copy phrase under each category, and an example of the formula is as follows:
[0124]
[0125] Where k represents the category to which the copy phrase belongs, and j represents any category in all categories.
[0126] After conversion, each copy phrase under each category can obtain its corresponding category dimension score ScoreCategory w .
[0127] It is not difficult to understand that the category dimension score is converted from the occurrence proportion of each copy phrase in the advertising copy of the same category, which quantifies the information contribution value of each copy phrase in the advertising copy of the same category from the perspective of category.
[0128] In an embodiment, when it is necessary to determine the store dimension score corresponding to the store dimension of a copy phrase, the following process can be referred to for implementation:
[0129] First, take the store as a unit, respectively count the word frequency of each copy phrase in the advertising copy of each store that has been launched under each category in the advertising copy of the advertising that has been launched by the store;
[0130] Each store contains multiple categories of corresponding goods, so it is possible to place advertisements for different categories of goods, so each category will likely contain multiple advertising copy, accordingly, based on the store, the number of occurrences of each phrase used by the store in the advertising copy of multiple same-category goods that have been placed in the store can be statistically determined, that is, its frequency freqency w_j_s It can be seen that the frequency here not only relates to the category to which the phrase library where the copy phrase is located belongs, but also relates to the advertising copy originating from the store, which is obtained by combining the two statistics.
[0131] Then, for each copy phrase, determine the store corresponding to the frequency higher than the preset threshold as the used store, and determine the total number of same-category stores that have placed advertisements for each category of goods and the total number of used stores.
[0132] For each copy phrase, if a store uses it less than a certain degree, the store's reference degree to the copy phrase is relatively weak, so a preset threshold can be used to determine the store that frequently uses the copy phrase. The preset threshold can be an empirical threshold or a measured threshold, which can be set by a person skilled in the art as needed. Specifically, for each store, compare the frequency of each copy phrase used by the store with the preset threshold. When the frequency is higher than the preset threshold, the store is determined to be a used store that frequently uses the copy phrase. For the case where the frequency is not higher than the preset threshold, the store can be determined to be a non-used store that uses the copy phrase less frequently.
[0133] For each copy phrase under each category, the corresponding used store can be determined according to the above principle, so the total number of used stores Store used_j In addition, for all stores that have placed advertisements for the same category j of goods, they can be determined as same-category stores that have placed advertisements for the category of goods, and then the total number of same-category stores Store all_j .
[0134] Further, multiply the frequency of each copy phrase under each store by the ratio of the total number of same-category stores to the total number of used stores to obtain the freshness of the copy phrase in the store dimension;
[0135] For each store, a copy phrase used by it in a category, when the total number of stores in the same category is given, if the total number of used stores using the copy phrase is higher, it means that the freshness is lower, on the contrary, the freshness is relatively high, and the role of the store for distinguishing other stores is higher, therefore, the degree of widespread use of each copy phrase can be determined by the ratio of the total number of stores in the same category to the total number of used stores, and further, the following formula can be used to determine the freshness ScoreStore of each copy phrase under each store and each category. w :
[0136]
[0137] Wherein, 1 is a regularization term for avoiding zero denominator, which can also be any very small number, and the word frequency freqency of the copy phrase w_j_s Here it can be regarded as an adjustment weight, it is not difficult to understand that the higher the word frequency, the higher the freshness of the copy phrase, which indicates that the store not only distinguishes other stores frequently using the copy phrase, but also is likely to be a common word distinguishing the store from other stores.
[0138] Finally, the freshness of all copy phrases corresponding to each store is normalized by category to obtain the store dimension score of each copy phrase under the corresponding store and category.
[0139] In order to facilitate the calculation of information contribution score, further, the maximum and minimum normalization processing method is applied to normalize the freshness of all copy phrases corresponding to each store by category, so as to obtain the store dimension score of each copy phrase under each store in each category. For the convenience of understanding, ScoreStore w represents the store dimension score.
[0140] It can be seen that through the above process, the store dimension score of the copy phrase in the phrase library of each category corresponding to each store using the copy phrase is determined, and based on the same set of advertising copy obtained from the advertising system, the store dimension score of the copy phrase used by each store is determined for each store. For each store, the store dimension score obtained for each copy phrase is personalized, which is related to the use frequency of the store to the copy phrase and the degree of widespread use of the copy phrase, has the freshness representation function, and quantifies the information contribution value of the copy phrase from the perspective of store use freshness.
[0141] In an embodiment, the advertising is taken as the dimension, the advertising dimension score of each copy phrase is determined according to the average effectiveness data of the effectiveness data obtained by each copy phrase in the advertising copy of the same category containing the copy phrase, and the advertising dimension score of each copy phrase is determined. The following process can be specifically implemented:
[0142] Firstly, determine the advertising copy of the same product category containing the advertising copy of each advertising phrase corresponding to the advertising phrase;
[0143] The phrase library can be used as a unit to determine the advertising copy corresponding to the category to which the phrase library belongs in the advertising copy set obtained by the advertising system. More specifically, for an advertising phrase in the phrase library, according to the category corresponding to the phrase library, the advertising copy using the advertising phrase in the advertising copy of the same product category is determined. These advertising copies are the advertising copies of the same product category corresponding to the advertising phrase.
[0144] Then, the effectiveness data of the advertising copy of the same product category corresponding to each advertising phrase is called from the advertising system;
[0145] For the advertising copy using the advertising phrase determined for each advertising phrase, the corresponding effectiveness data can be further called from the advertising system.
[0146] Further, the average effectiveness data of each advertising phrase in the corresponding category is obtained by averaging the effectiveness data of the advertising copy of the same product category of each advertising phrase;
[0147] For each advertising phrase, the average effectiveness data of the advertising copy of the same product category is obtained by averaging the effectiveness data of the advertising copy of the same product category, and the average click-through rate CTR corresponding to each advertising phrase is obtained. aveage .
[0148] Finally, the average effectiveness data of each advertising phrase is normalized by category to obtain the advertising dimension score of each advertising phrase in each category.
[0149] In order to facilitate the calculation of information contribution score, further, the maximum and minimum normalization processing method is applied to normalize the average effectiveness data of each advertising phrase by category, and the advertising dimension score ScoreCTR of each advertising phrase in each category is obtained. w .
[0150] As can be seen, through the above process, the category to which the advertising phrase belongs is associated, and the advertising dimension score corresponding to each advertising phrase is quantified according to the effectiveness data of the advertising copy of the same product category. The advertising dimension score has the function of representing the advertising effectiveness obtained by the advertising copy using the advertising phrase, and quantifies the information contribution value of the advertising phrase from the perspective of advertising effectiveness.
[0151] In one embodiment, in order to realize the synthesis of the scores obtained in each different dimension, each score corresponding to each copy phrase is weighted and summarized according to each copy phrase in the phrase library corresponding to each category in each store, to obtain the information contribution score of each copy phrase under different stores w The exemplary formula is as follows:
[0152] Score w = c1*ScoreCategory w + c2*ScoreStore w + c3*ScoreCTR w
[0153] Wherein, c1, c2, c3 are respectively preset weights corresponding to the category dimension score, the store dimension score, and the advertisement dimension score of the copy phrase, which can be preset by the person skilled in the art as needed.
[0154] According to the formula, it is not difficult to understand that the store dimension score is introduced in the determination of the information contribution score, and the store dimension score is determined for each copy phrase in each store. Therefore, the information contribution score obtained according to the formula is actually determined according to each store. Similarly, the category dimension score is also included in the formula, and the copy phrase itself may belong to different categories, so the information contribution score also needs to be associated with the category used by the copy phrase to determine. As can be seen, the information contribution score is determined under the condition of a specified store and a specified category. When calculating the information contribution score of a copy phrase, the category and the store need to be given as constraint conditions, so as to call the corresponding category dimension score and store dimension score, and weightedly summarize the advertisement dimension score to obtain the final information contribution score.
[0155] According to the above principle, each store actually has a phrase library corresponding to each category, and each copy phrase in each phrase library stores the information contribution score corresponding to each category under the condition of the store.
[0156] As can be understood from the above embodiment, the present embodiment uses a standardized processing process based on the copy phrases extracted from the advertisement copy of the advertisement system, which not only quantifies the information contribution value of each copy phrase from the category dimension, but also quantifies the information contribution value of each copy phrase corresponding to each store dimension.
[0157] According to the extended embodiment of any of the above embodiments, please refer to Figure 4 The step S1300 of determining the similarity and confidence between the query statement and each detail statement matched out comprises:
[0158] Step S1310, the query sentence and each matched detail sentence form a sentence pair, and the text matching classification model pre-trained to a convergent state is input to synchronously determine the classification probabilities of each category of different matching degrees of the first classification space corresponding to the sentence pair and the classification probability of the category of the second classification space corresponding to whether the sentence pair is suitable for generalization.
[0159] In this embodiment, a preset text matching classification model is used to determine the similarity and confidence corresponding to the detail sentence recalled according to the script phrase. An exemplary text matching classification model is shown in Figure 5 The text feature extraction model is preferably a Bert model. In one embodiment, the query sentence and the detail sentence can be constructed as the input of the model by adding the [SEP] label in the Bert model and prepositioning [CLS], and instructing the model to perform the next sentence recognition task, in the form of:
[0160] [CLS] query sentence [SEP] detail sentence
[0161] The text matching classification model is trained to a convergent state in advance using training samples in a corresponding preset data set, so that it learns the ability to synchronously determine the similarity and confidence corresponding to the input sentence pair.
[0162] In the text matching classification model, the first classifier can be a multi-classifier. The first classification space corresponding to the first classifier can be set to correspond to multiple categories according to the matching degree of the query sentence and the detail sentence in the sentence pair, each category corresponding to a semantic closeness level between the query sentence and the detail sentence, and the classification probability obtained by classifying each category can be used as a representation of the similarity between the query sentence and the detail sentence.
[0163] For example, the categories of the second classification space can be set to three, representing that the query sentence and the detail sentence are completely irrelevant, partially relevant, and closely related, respectively, thereby establishing a classification scoring standard for the matching degree LabelMatch between the detail sentence and the product title:
[0164] LabelMatch = 0, representing the first category of complete irrelevance, indicating that the detail sentence cannot effectively represent the product characteristics of the advertised product, such as product after-sales, discount promotion, working principle, product maintenance, logistics transportation, etc.
[0165] LabelMatch = 1, representing the second category of partial correlation, the first detail sentence can effectively represent the product characteristics, then, satisfying condition 1: the detail sentence represents the basic function of the product but is not the core selling point; and / or, satisfying condition 2: the expression of the detail sentence and the product features of the advertised product are partially matched.
[0166] LabelMatch = 2, representing the third category of close correlation, the detail sentence completely matches the product features and can very appropriately and sufficiently express the core selling point of the product.
[0167] In the text matching classification model, the second classifier can be a binary classifier. The second classification space corresponding to the second classifier can set two categories corresponding to whether the query sentence and the detail sentence in the sentence pair are suitable for advertising promotion, and the classification probability obtained by classifying each category can be used as the confidence between the corresponding whether it is suitable for promotion.
[0168] For example, the categories of the second classification space can be set to two, respectively representing the query sentence and the detail sentence, mainly the detail sentence, suitable or not suitable for being used as an element of the advertising copy, that is, whether the corresponding detail sentence is suitable for promotion. Accordingly, a classification scoring standard of the detail sentence for the promoteability LabelPromote of the product can be established:
[0169] LabelPromote = 1, indicating that the detail sentence is very suitable as a marketing promotion copy, which introduces the functional selling point of the product from one or more of the following three aspects: the revenue that the product can bring, the use scenario of the product, and the detailed parameter description of the function of the product.
[0170] LabelPromote = 0, not satisfying the statement of LabelPromote = 1.
[0171] It should be noted that the product features mainly refer to the selling point characteristics of the product, that is, the product description content with information reference value that attracts consumers to purchase the corresponding product.
[0172] Step S1320, determining the classification probability of the matching category of the detail sentence in the sentence pair in the first classification space as the similarity of the detail sentence in the sentence pair to the matching category;
[0173] According to the foregoing, the text matching classification model is provided with multiple categories, when the comprehensive feature vector of a sentence pair is processed by the first classifier to obtain the classification probability mapped to each category, the category with the largest classification probability is the corresponding category of the sentence pair, and the classification probability of the category can be used as the similarity of the detail sentence in the sentence pair.
[0174] Step S1330, determining the confidence of the detail sentence in the sentence pair corresponding to the category suitable for promotion as the classification probability of the category suitable for promotion of the detail sentence in the second classification space.
[0175] According to the foregoing, when the comprehensive feature vector of a sentence pair is processed by the second classifier to obtain the classification probability of each category to which it is mapped, only the category corresponding to the positive result, such as the classification probability corresponding to LabelPromote=1 in the foregoing example, can be taken as the corresponding confidence.
[0176] Step S1340, establishing a mapping relationship between each detail sentence and the similarity of the category corresponding to the first classification space and the confidence of the category corresponding to the second classification space.
[0177] By inputting each sentence pair into the text matching classification model one by one, the similarity and confidence corresponding to each detail sentence can be determined synchronously, the mapping relationship between each detail sentence and its similarity and confidence is established, and the mapping relationship data is determined, which can be directly called later.
[0178] According to the above embodiments, it can be seen that the similarity and confidence corresponding to each detail sentence can be determined synchronously by means of the same neural network model. The similarity can be used to represent whether the detail sentence is closely related to the selling point of the product title in the query sentence, and the confidence can be used to represent whether the detail sentence is suitable for promotion. It is not difficult to understand that according to the combination of the two, the information contribution price of each detail sentence for creating an advertising script can be more effectively represented. Later, the detail sentences can be optimized according to the two, and the advertising product can accurately match the detail sentences that can better reflect the promotion value from the product detail text.
[0179] According to the embodiments extended according to any of the above embodiments, please refer to Figure 6 , the training process of the text matching classification model comprises:
[0180] Step S3100, calling a single training sample in a preset data set into the text matching classification model, each training sample being associated with a first label and a second label, and comprising a sample query sentence and a sample detail sentence. The sample query sentence comprises a product title and a category label of a historical advertising product, and the sample detail sentence is a detail sentence extracted from the detail text of the historical advertising product. The first label is used to indicate the category corresponding to the matching degree between the sample query sentence and the sample detail sentence, and the second label is used to indicate the category corresponding to whether the sample detail sentence is suitable for promotion.
[0181] Please continue to refer toFigure 5 The text matching classification model shown, as described in the foregoing, is correspondingly provided with two branches so as to map the comprehensive features obtained by the text feature extraction model thereof to two classifiers to determine the similarity and the confidence. The two classifiers correspondingly have a first classification space and a second classification space, which, according to the example of the previous embodiment, are respectively provided with multiple categories and two categories.
[0182] In order to train such a text matching classification model to effectively obtain the similarity and the confidence corresponding to each detail sentence, in this embodiment, a data set is prepared for iterative training, and the training is performed to convergence to obtain the corresponding ability.
[0183] The data set includes training samples sufficient to train the text matching classification model to convergence, and the training samples are constructed according to the principles disclosed in the foregoing, that is, each training sample includes a sample query sentence composed of a product title and a category label of a product, and a sample detail sentence extracted from a detail sentence of the product. The sample query sentence and the sample detail sentence form a sentence pair, that is, a training sample is constructed. In one embodiment, the information used to construct the training sample belongs to a product that has placed an advertisement in an advertisement system, that is, a historical advertisement product.
[0184] In order to supervise the training process, two supervision labels are provided for the outputs of the two classifiers corresponding to each training sample. The supervision labels can be manually labeled labels, which are a first label and a second label respectively. The first label is provided for the input of the first classifier to supervise the first classifier, which indicates the category to which the detail sentence in the training sample should be matched in the first classification space. The second label is provided for the input of the second classifier to supervise the second classifier, which indicates the category to which the detail sentence in the training sample should be matched in the second classification space. In one embodiment, when manually labeling, the labeler can subjectively evaluate the information contribution value of the detail sentence to determine the corresponding first label and second label. In another embodiment, the first label can be determined according to the similarity between the feature vectors of the corresponding detail sentence and the product title in the query sentence, and the second label can be determined according to the advantages and disadvantages of the advertisement effectiveness data obtained by the advertisement of the corresponding historical advertisement product.
[0185] The category structure of the first classification space and the second classification space can be pre-set before training.
[0186] Step S3200, deep semantic information of the training sample is extracted by the text matching classification model, and two-way classification mapping is synchronously performed according to the deep semantic information, and is respectively mapped to the first classification space and the second classification space, to obtain classification probabilities of respective categories in the first classification space and the second classification space, and to determine target categories corresponding to the training sample in the first classification space and the second classification space according to the classification probabilities;
[0187] As described above, the text matching classification model extracts deep semantic information of an embedding vector of a training sample input therein through a text feature extraction model inside the text matching classification model, obtains a corresponding comprehensive feature vector, and then the comprehensive feature vector enters a branch corresponding to a first classifier and a second classifier, respectively, and is classified and mapped to a corresponding output layer in each classifier through an internal full connection layer corresponding to each classification space, and the output layer calculates classification probabilities mapped to respective categories according to a preset classification function. Thus, the first classification space and the second classification space can obtain classification probabilities corresponding to respective categories.
[0188] For the first classification space, the category with the maximum classification probability is the category corresponding to the matching degree obtained by the training sample, and the classification probability can be used as a similarity. For the second classification space, the category with the maximum classification probability is also the category corresponding to whether the training sample is suitable for promotion, and the classification probability corresponding to the category suitable for promotion can be used as a confidence in the subsequent model inference stage.
[0189] Step S3300, a first loss value obtained by calculating a loss of a target category of the first classification space according to a first label, a second loss value obtained by calculating a loss of a target category of the second classification space according to a second label, and a model loss value obtained by summarizing the first loss value and the second loss value;
[0190] In order to make the text matching classification model converge, it is necessary to calculate the loss value for the classification results of the first classification space and the second classification space. Specifically, a first loss value corresponding to a target category in the first classification space is calculated according to a first label corresponding to the training sample, and a second loss value corresponding to a target category in the second classification space is calculated according to a second label corresponding to the training sample. As can be seen, the two labels are used to calculate the classification results of the two classifiers, respectively, and are supervised respectively.
[0191] In order to realize joint training, on the basis of obtaining the first loss value and the second loss value, the first loss value and the second loss value are further matched with preset weights for summation, that is, the weighted sum value of the two is calculated to realize the summary of the loss value, and the summarized model loss value is obtained. It is not difficult to understand that by pursuing the minimization of the model loss value, the text matching classification model can be continuously trained to a convergent state.
[0192] Step S3400, judging whether the text matching classification model converges according to the model loss value, when not converging, implementing gradient update on the text matching classification model, and continuing to call the next training sample for iterative training.
[0193] In order to control the training process of the text matching classification model, a preset threshold is provided, and the model loss value corresponding to each training sample is compared with the preset threshold. When the model loss value reaches the preset threshold, it indicates that the model has converged, and the iterative training of the model can be terminated, and the model can be put into the inference stage for use. When the model loss value does not reach the preset threshold, it indicates that the model has not converged, and therefore, the model loss value can be used for back propagation to realize gradient update of the weights of each link of the model, so that the model is further approximated to converge, and then the next training sample in the data set is called to implement the next iteration of the model. training. By analogy, until the model is trained to the convergent state.
[0194] According to the above embodiments, by providing the first label and the second label corresponding to the similarity and the confidence of the training sample in the training stage, the text matching classification model is implemented for joint training, the loss values corresponding to the similarity and the confidence are calculated respectively in the training process, and the model loss value is finally calculated and summarized to pursue the minimization of the model loss value. Thus, the model can simultaneously determine the corresponding similarity and confidence of the sentence pair, so that the model has the function of giving the information contribution value quantization data of the detail sentence in the sentence pair for the advertising copywriting of the advertising product. In the embodiment, the training samples in the data set can be obtained from the e-commerce platform and the advertising system, which is easy to prepare in batches and has high training efficiency.
[0195] According to the extended embodiments of any of the above embodiments, please refer to Figure 7 , the step S1400, screening part of the detail sentences according to the similarity and the confidence, comprising:
[0196] Step S1410, taking the category of the first classification space as the main index, and performing the first reverse sorting on each detail sentence according to the similarity;
[0197] Each category in the first classification space itself represents the semantic closeness of the detail sentence and the advertising product in terms of product characteristics, so that the category label itself can be used to quickly filter the detail sentences. In the process of screening the detail sentences, the category label LabelMatch obtained by each detail sentence in the first classification space can be used as the main index to perform the first reverse sorting on each detail sentence, and a first list is obtained.
[0198] Step S1420, performing second reverse sorting on each detail statement in the first sorted list according to a weighted sum value of the similarity and the confidence of each matched detail statement;
[0199] The similarity and the confidence of each detail statement can comprehensively represent the information contribution value of the detail statement, and thus, the corresponding weighted sum value can be obtained by weighting and summing the similarity and the confidence corresponding to each detail statement as a comprehensive score of each detail statement, and then the second reverse sorting is performed on the detail statements in the first list according to the comprehensive score as a secondary index to obtain a second list. The weights corresponding to the similarity and the confidence can be set as needed. In one embodiment, the weights can be the weights configured by the text matching classification model for the first loss value and the second loss value when calculating the total loss value of the model in the training stage, so as to maintain the consistency of the information value measurement standard as much as possible.
[0200] Step S1430, selecting a predetermined number of detail statements ranked in the front from the second sorted list as the copy material of the advertisement copy of the advertisement product.
[0201] As can be easily understood, in the second list obtained by the above two reverse sortings, the quality of the part of the detail statements is relatively high, and the precision sorting of the detail statements matched from the detail text of the advertisement product by using the copy phrase has been realized. Therefore, as long as a predetermined number of detail statements ranked in the front are selected from the second list, these detail statements are the text material suitable for making the advertisement copy.
[0202] As can be seen, in the above embodiments, the categories of the first classification space play a role in rough sorting of the detail text, and the importance of whether the detail statement represents the product features of the advertisement product is highlighted. On this basis, the weighted sum value of the similarity and the confidence is used for precision sorting, and finally the sorted and selected detail statements are used as the copy material. The closeness of the detail statement to the product features of the advertisement product is considered first, and on the basis of the same closeness, the degree of the detail statement suitable for promotion is considered first. The copy material determined in this way has reasonable information reference value sorting and high reference efficiency, and can improve the efficiency and accuracy of the copy material calling of the advertisement copy creation.
[0203] According to the embodiments extended according to any of the above embodiments, please refer to Figure 8 , after the copy material list is constituted, the method further comprises:
[0204] Step S1500, pushing the copy material list to a terminal device displaying the advertisement product;
[0205] After the copywriting materials matched from the detail text of the advertised commodity are structured as a copywriting material list, the list can be pushed to a terminal device that submitted the advertised commodity for display, that is, in response to an event in which a user invokes an advertisement publishing interface, the user's specified copywriting material list corresponding to the advertised commodity is pushed to the user, and the copywriting materials in the list are selected from the detail text of the advertised commodity, facilitating the user to quickly extract the sentences that represent the features of the advertised commodity.
[0206] Step S1600, in response to the advertisement publishing request submitted by the terminal device, obtaining the corresponding advertisement copy, wherein the advertisement copy contains the copywriting materials referenced from the copywriting material list;
[0207] When the user references the copywriting materials in the copywriting material list on the terminal device and completes the configuration of the advertisement, the user can submit an advertisement publishing request. In response to the advertisement publishing request, the server can obtain the advertisement copy finally submitted by the user through the request. Generally, the user will reference the copywriting materials, and therefore, the advertisement copy generally contains one or more copywriting materials in the copywriting material list. Of course, the user can also make appropriate editing on the basis of the copywriting materials, so as to represent the revised and updated versions of the copywriting materials.
[0208] Step S1700, publishing the advertisement corresponding to the advertised commodity with the advertisement copy.
[0209] After the server obtains the advertisement copy, the advertisement copy is published to the advertisement system, the advertisement corresponding to the advertised commodity is published through the advertisement system, and the terminal device of the user receiving the advertisement can read the advertisement copy containing the corresponding copywriting materials.
[0210] According to the above embodiments, it can be known that the present application can provide the user with the detail sentences with information contribution value as copywriting materials in the process of writing the advertisement copy by the user, which has the function of creative inspiration, can improve the self-creation efficiency of the advertisement copy, and can guide the user to output high-quality advertisement copy.
[0211] Please refer to Figure 9To adapt to one of the purposes of the present application, a text material extraction device is provided, which is a functional embodiment of the text material extraction method of the present application. The device comprises a query construction module 1100, a sentence recall module 1200, a matching processing module 1300, and a material generation module 1400. The query construction module 1100 is used to obtain the title text and category label of an advertising product, and construct a query sentence. The sentence recall module 1200 is used to recall detail sentences in the detail text of the advertising product according to the text phrases matched with the title text and / or category label. The matching processing module 1300 is used to determine the similarity and confidence between the query sentence and each matched detail sentence. The material generation module 1400 is used to filter part of the detail sentences as the text material of the advertising text of the advertising product according to the similarity and confidence, and constitute a text material list.
[0212] In an embodiment extended according to any of the above embodiments, the sentence recall module 1200 comprises a detail sentence unit for dividing the detail text of the advertising product into sentences to obtain a sentence list composed of various detail sentences in the detail text; a phrase matching unit for matching a plurality of text phrases from a phrase library corresponding to the category label according to the title text and / or category label, to constitute a phrase list, wherein the text phrases comprise a plurality of word units with independent word features; and a sentence screening unit for calculating the similarity between each text phrase in the phrase list and each detail sentence in the sentence list, and screening the detail sentences semantically matched with each text phrase according to the similarity.
[0213] In an embodiment extended according to any of the above embodiments, before the text phrases matched with the title text and / or category label, a phrase extraction unit is used to extract a plurality of text phrases from the advertising text of the already launched advertising corresponding to the category label in the advertising system to constitute candidate phrases, wherein the text phrases are extracted according to a plurality of preset phrase structures, and the phrase structure comprises a plurality of ordered word feature labels, at least one of which is a word feature label representing a noun, which is postfixed relative to other word feature labels; a phrase scoring unit is used to determine the information contribution score of each candidate phrase by referring to the category, store, and advertising of the candidate phrase; and a screening and library building unit is used to screen part of the candidate phrases according to the information contribution score, and reserve the text phrases stored as the phrase library.
[0214] In an embodiment extended from any of the above embodiments, the matching processing module 1300 comprises: a sentence pair input unit configured to input a query sentence and each matched detail sentence into the pre-trained text matching classification model in a converged state to simultaneously determine classification probabilities of each category of different matching degrees represented by a first classification space and classification probabilities of a category of whether the detail sentence is suitable for promotion represented by a second classification space corresponding to the sentence pair; a first classification unit configured to determine a similarity of the detail sentence in the sentence pair corresponding to the matched category by taking the classification probability of the category of matching the query sentence and the detail sentence in the sentence pair as represented by the first classification space as a criterion; a second classification unit configured to determine a confidence of the detail sentence in the sentence pair corresponding to the category of being suitable for promotion by taking the classification probability of the category of the detail sentence being suitable for promotion as represented by the second classification space as a criterion; and a data storage unit configured to establish a mapping relationship between the similarity of the category corresponding to the first classification space and the confidence of the category corresponding to the second classification space of each matched detail sentence.
[0215] In an embodiment extended from any of the above embodiments, the script material extraction device comprises a training module configured to perform a training process of the text matching classification model, comprising: a sample calling unit configured to input a single training sample in a preset data set into the text matching classification model, each training sample being associated with a first label and a second label and comprising a sample query sentence and a sample detail sentence, the sample query sentence comprising a product title and a category label of a historical advertising product, and the sample detail sentence being a detail sentence extracted from a detail text of the historical advertising product, the first label being used to indicate categories corresponding to a plurality of matching degrees between the sample query sentence and the sample detail sentence, and the second label being used to indicate whether the sample detail sentence is suitable for promotion of a corresponding category; a training execution unit configured to extract deep semantic information of the training sample by the text matching classification model, and simultaneously perform two-way classification mapping according to the deep semantic information to the first classification space and the second classification space respectively to obtain classification probabilities of each category in the first classification space and the second classification space, and determine target categories of the training sample in the first classification space and the second classification space according to the classification probabilities; a loss calculation unit configured to calculate a first loss value obtained by calculating a loss of the target category in the first classification space according to the first label, calculate a second loss value obtained by calculating a loss of the target category in the second classification space according to the second label, and aggregate the first loss value and the second loss value into a model loss value; and an iteration decision unit configured to determine whether the text matching classification model converges according to the model loss value, and when the text matching classification model does not converge, perform gradient update on the text matching classification model and continue to call a next training sample for iteration training.
[0216] In an embodiment expanded according to any of the above embodiments, the material generation module 1400 comprises: a first sorting unit configured to sort the matched detail statements according to the similarity in a first reverse order with the category of the first classification space as the main index; a second sorting unit configured to sort the detail statements after the first sorting in a second reverse order according to the weighted sum of the similarity and the confidence of each detail statement; and a material selection unit configured to select a predetermined number of detail statements with high rankings from the detail statements after the second sorting as the copywriting materials of the advertising copy of the advertising product.
[0217] In an embodiment expanded according to any of the above embodiments, the material generation module 1400 comprises: a list pushing unit configured to push the list of copywriting materials to a terminal device that submits the advertising product for display; a copywriting acquisition unit configured to acquire a corresponding advertising copy in response to an advertising publishing request submitted by the terminal device, wherein the advertising copy contains copywriting materials cited from the list of copywriting materials; and an advertising publishing unit configured to publish an advertisement corresponding to the advertising product with the advertising copy.
[0218] To solve the above technical problems, the embodiments of the present application further provide a computer device. As shown in Figure 10 The computer device includes a processor, a computer readable storage medium, a memory and a network interface connected through a system bus. The computer readable storage medium of the computer device stores an operating system, a database and computer readable instructions. The database can store a control information sequence. The computer readable instructions, when executed by the processor, can enable the processor to implement a product search category identification method. The processor of the computer device is configured to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device can store computer readable instructions. The computer readable instructions, when executed by the processor, can enable the processor to execute the copywriting material extraction method of the present application. The network interface of the computer device is configured to connect and communicate with a terminal. Those skilled in the art can understand that the structure shown in Figure 10 is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0219] In the present embodiment, the processor is configured to execute Figure 9The specific functions of each module and its sub-modules in the above embodiment are described above. The memory stores the program codes and various data required for executing the above modules or sub-modules. The network interface is used for data transmission between the user terminal and the server. The memory in the above embodiment stores the program codes and data required for executing all modules / sub-modules of the script material extraction device of the present application, and the server can call the program codes and data of the server to execute the functions of all sub-modules.
[0220] The present application also provides a storage medium storing computer readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps of the script material extraction method of any embodiment of the present application.
[0221] The present application also provides a computer program product comprising computer programs / instructions, which, when executed by one or more processors, implement the steps of the method described in any embodiment of the present application.
[0222] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiments of the present application can be completed by a computer program instructing related hardware, and the computer program can be stored in a computer readable storage medium. The program, when executed, can include the processes of the above-mentioned embodiments of the method. The storage medium can be a computer readable storage medium such as a magnetic disc, an optical disc, a read-only memory (ROM), or a random access memory (RAM).
[0223] In summary, the present application can extract high-quality detail sentences that can describe the characteristics of the product from the detail text of the advertisement product to be published as script materials for users to quote. The present application can realize the auxiliary creation of the advertisement script. For the process of writing the advertisement script, the user can quote the script materials as needed.
[0224] A person of ordinary skill in the art can understand that the steps, measures, and schemes in the various operations, methods, and processes discussed in the present application can be alternated, changed, combined, or deleted. Further, other steps, measures, and schemes in the various operations, methods, and processes discussed in the present application can also be alternated, changed, rearranged, decomposed, combined, or deleted. Further, the steps, measures, and schemes in the prior art with the various operations, methods, and processes disclosed in the present application can also be alternated, changed, rearranged, decomposed, combined, or deleted.
[0225] The above merely describes some embodiments of the present application, and it should be pointed out that, for those skilled in the art, some improvements and refinements can be made without departing from the principles of the present application, and these improvements and refinements should also be considered as the protection scope of the present application.
Claims
1. A method of extracting a script material, characterized by, The method comprises the following steps: Obtain the title text and category label of the advertised commodity, and construct a query statement; According to the text phrases matched with the title text and / or category label, recall the detail sentences in the detail text of the advertised commodity matched with the text phrases, and the text phrases are extracted from the advertised text of the advertised commodity corresponding to the category label according to a preset plurality of phrase structures, the phrase structure comprises a plurality of ordered part-of-speech tags, at least one of which is a noun tag, which is postfixed relative to other part-of-speech tags; Determine the similarity and confidence between the query statement and each matched detail sentence, and the confidence is used to represent whether the detail sentence is suitable for promotion; According to the similarity and confidence, filter out part of the detail sentences as the text materials of the advertised text of the advertised commodity, and construct a text material list.
2. The script material extraction method according to claim 1, characterized by, According to the text phrases matched with the title text and / or category label, recall the detail sentences in the detail text of the advertised commodity matched with the text phrases, comprising: Divide the detail text of the advertised commodity into sentences to obtain a sentence list composed of various detail sentences in the detail text; According to the title text and / or category label, match a plurality of text phrases from the phrase library corresponding to the category label to form a phrase list, and the text phrases comprise a plurality of word units with independent word features; Calculate the similarity of each text phrase in the phrase list with each detail sentence in the sentence list, and filter out the detail sentences that are semantically matched with each text phrase according to the similarity.
3. The script material extraction method according to claim 1, characterized by, Before the text phrases matched with the title text and / or category label, comprising: Extract a plurality of text phrases from the advertised text of the advertised commodity corresponding to the category label in the advertising system to form a candidate phrase; Determine the information contribution score of each candidate phrase by referring to the category, store, and advertisement of the candidate phrase; According to the information contribution score, filter out part of the candidate phrases, and retain the text phrases stored in the phrase library.
4. The script material extraction method according to claim 1, characterized by, Determine the similarity and confidence between the query statement and each matched detail sentence, comprising: Input the query statement and each matched detail sentence into a text matching classification model pre-trained to a convergent state to simultaneously determine the classification probability of each category representing different matching degrees in a first classification space and the classification probability of a category representing whether it is suitable for promotion corresponding to the sentence pair; Determine the similarity of the detail sentence in the sentence pair corresponding to the matched category as the classification probability of the category representing the matching of the query statement and the detail statement in the sentence pair in the first classification space; Determine the confidence of the detail sentence in the sentence pair corresponding to the category suitable for promotion as the classification probability of the category representing the suitability for promotion of the detail statement in the sentence pair in the second classification space; Establish a mapping relationship between the similarity of each matched detail sentence corresponding to the category in the first classification space and the confidence of the category in the second classification space.
5. The script material extraction method according to claim 4, characterized by, The training process of the text matching classification model comprises: The preset dataset is called to input a text matching classification model, each training sample is associated with a first label and a second label, and includes a sample query statement and a sample detail statement, the sample query statement includes a product title and a category label of a historical advertising product, and the sample detail statement is a detail statement extracted from a detail text of the historical advertising product, the first label is used to indicate a category corresponding to a plurality of matching degrees between the sample query statement and the sample detail statement, and the second label is used to indicate whether the sample detail statement is suitable for promoting a corresponding category; Deep semantic information of the training sample is extracted by the text matching classification model, and two-way classification mapping is synchronously performed according to the deep semantic information, and is mapped to a first classification space and a second classification space respectively, to obtain classification probabilities of each category in the first classification space and the second classification space, and to determine target categories of the training sample in the first classification space and the second classification space according to the classification probabilities; A first loss value is obtained by calculating a loss of the target category of the first classification space according to the first label, a second loss value is obtained by calculating a loss of the target category of the second classification space according to the second label, and the first loss value and the second loss value are summarized as a model loss value; Whether the text matching classification model converges is determined according to the model loss value, when the text matching classification model does not converge, gradient updating is performed on the text matching classification model, and the next training sample is iteratively trained.
6. The script material extraction method according to claim 4, characterized by, Part of the detail statements are screened according to the similarity and the confidence, including: The first classification space is taken as a main index, and each detail statement matched is first sorted in reverse according to the similarity; The weighted sum of the similarity and the confidence of each detail statement matched is taken as a second sorting index, and each detail statement after the first sorting is second sorted in reverse; A predetermined number of detail statements ranked in the front are selected from each detail statement after the second sorting as copywriting materials of an advertising copy of an advertising product.
7. The script material extraction method according to any one of claims 1 to 6, characterized by, After the copywriting material list is formed, including: The copywriting material list is pushed to a terminal device for displaying; According to an advertising publishing request submitted by the terminal device, a corresponding advertising copy is obtained, the advertising copy includes copywriting materials cited from the copywriting material list; The advertising product corresponding to the advertising copy is published.
8. An article extraction apparatus, characterized by comprising: Including: A query construction module is configured to obtain a title text and a category label of an advertising product, and construct a query statement; A statement recall module is configured to recall detail statements in detail texts of an advertising product according to copywriting phrases matched with the title text and / or the category label, the copywriting phrases are extracted from advertising copies of already launched advertising products corresponding to the category label according to a preset plurality of phrase structures, and the phrase structure includes a plurality of ordered word tags, at least one of which is a word tag representing a noun, and the word tag is postfixed relative to other word tags. The matching processing module is configured to determine similarity and confidence between the query statement and each matched detail statement, the similarity and the confidence being determined synchronously for each detail statement, and the confidence being used to represent whether the detail statement is suitable for promotion. The material generation module is configured to filter part of the detail statements as copy materials of the advertisement copy of the advertisement commodity according to the similarity and the confidence, and form a copy material list.
9. A computer device comprising a central processing unit and a memory, characterized in that The central processing unit is configured to call and run a computer program stored in the memory to perform the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer program is stored in the form of computer readable instructions and is implemented according to the method of any one of claims 1 to 7, and when the computer program is called and run by a computer, the steps included in the corresponding method are performed.
Citation Information
Patent Citations
Social advertising facing Twitter feasibility analysis method
CN104268130A
Intelligent recommendation method and system for network materials, computer equipment and medium
CN113849462A
Commodity information matching method and device, commodity information publishing method and device, equipment, medium and product
CN114065750A