A data processing method and apparatus based on marketing text
By tagging and attribute-processing marketing texts to generate the first triplet information, the problem of low efficiency in marketing text tagging processing is solved, the quality of key information extraction is improved, and an effective knowledge graph is constructed.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-08
- Publication Date
- 2026-03-10
AI Technical Summary
The existing marketing text tagging process is inefficient, resulting in poor quality of key information extraction and difficulty in building an effective knowledge graph.
Marketing text information is classified and identified using tag processing rules and attribute processing rules to obtain tag information and attribute information, and the first triplet information is constructed to generate a knowledge graph.
It improved the efficiency of tagging marketing texts, enhanced the quality of key information extraction, and constructed a more effective knowledge graph.
Smart Images

Figure CN114357157B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a data processing method and apparatus based on marketing text. Background Technology
[0002] Current text generation models commonly use open-source tools and train model parameters, then fine-tuning them with domain-specific datasets. However, the relatively low efficiency of tagging marketing text results in extracted tags lacking logic and coherence. Therefore, providing a data processing method and apparatus based on marketing text to improve its tagging efficiency, thereby enhancing the quality of key information extraction and constructing an effective knowledge graph, is of paramount importance. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a data processing method and apparatus based on marketing text, which can obtain tag information and attribute information by processing marketing text information to be processed through tag processing rules and attribute processing rules, and then perform comprehensive processing to obtain the first triplet information for constructing a knowledge graph that guides the generation of marketing text. This is beneficial to improving the tagging efficiency of marketing text, thereby improving the quality of key information extraction, so as to construct an effective knowledge graph.
[0004] To address the aforementioned technical problems, a first aspect of the present invention discloses a data processing method based on marketing text, the method comprising:
[0005] Obtain marketing text information to be processed;
[0006] The marketing text information to be processed is classified and processed using preset tag processing rules to obtain tag information; the tag information is related to the industry and text style of the marketing text information to be processed.
[0007] The marketing text information to be processed is identified and processed using preset attribute processing rules to obtain attribute information;
[0008] The tag information and the attribute information are processed to obtain the first triplet information; the first triplet information includes M element information; M is a positive integer greater than or equal to 3; the element information is related to the tag information and the attribute information; the first triplet information is used to construct a knowledge graph for guiding the generation of marketing text.
[0009] As an optional implementation, in the first aspect of the present invention, the step of classifying the marketing text information to be processed using preset tag processing rules to obtain tag information includes:
[0010] The marketing text information to be processed is cleaned using preset text cleaning rules to obtain text information to be segmented; the cleaning process is used to filter out non-standard information.
[0011] The text information to be segmented is processed to obtain text information to be classified; the segmentation process is used to reduce the uncertainty of the text information to be segmented.
[0012] The text information to be classified is labeled using a preset classification model to obtain tag information; the tag information includes industry information and text style information.
[0013] As an optional implementation, in the first aspect of the present invention, the step of using preset attribute processing rules to identify and process the marketing text information to be processed to obtain attribute information includes:
[0014] The marketing text information to be processed is processed using a preset recognition model to obtain text output information;
[0015] Determine whether the text output information meets the verification conditions, and obtain the judgment result;
[0016] When the judgment result is yes, the text output information is determined to be attribute information;
[0017] When the judgment result is negative, the marketing text information to be processed is processed using preset new word extraction rules to obtain the attribute information.
[0018] As an optional implementation, in the first aspect of the present invention, the step of processing the marketing text information to be processed using preset new word extraction rules to obtain the attribute information includes:
[0019] The marketing text information to be processed is processed by using a preset new word extraction model to obtain word fragment information;
[0020] The word fragment information is classified and filtered to obtain the attribute information.
[0021] As an optional implementation, in the first aspect of the present invention, the classification model includes a first classification model, and / or a second classification model, and / or a third classification model;
[0022] The step of labeling the text information to be classified using a preset classification model to obtain tag information includes:
[0023] The industry information is obtained by labeling the text information to be classified using the first classification model;
[0024] The text information to be classified is labeled using the second classification model to obtain the text style information; or,
[0025] The third classification model is used to label the text information to be classified, thereby obtaining the industry information and the text style information.
[0026] As an optional implementation, in the first aspect of the present invention, after the marketing text information to be processed is identified and processed using preset attribute processing rules to obtain attribute information, the method further includes:
[0027] The marketing text information to be processed is subjected to word extraction and filtering to obtain the word segmentation information to be aggregated;
[0028] The word segmentation information to be clustered is subjected to clustering processing to obtain cluster information; the cluster information includes cluster information and category name information;
[0029] The cluster information and the attribute information are processed to obtain the second triplet information.
[0030] As an optional implementation, in the first aspect of the present invention, the step of performing clustering processing on the word segmentation information to be clustered to obtain cluster information includes:
[0031] The word segmentation information to be clustered is processed using a preset clustering model to obtain the cluster information; the cluster information includes L clusters; where L is a positive integer greater than or equal to 3;
[0032] The cluster information is labeled to obtain the category name information; the category name information includes the L category names.
[0033] A second aspect of this invention discloses a data processing apparatus based on marketing text, the apparatus comprising:
[0034] The acquisition module is used to acquire marketing text information to be processed;
[0035] The first processing module is used to classify the marketing text information to be processed using preset tag processing rules to obtain tag information; the tag information is related to the industry and text style of the marketing text information to be processed.
[0036] The second processing module is used to identify and process the marketing text information to be processed using preset attribute processing rules to obtain attribute information.
[0037] The third processing module is used to process the tag information and the attribute information to obtain the first triplet information; the first triplet information includes M element information; M is a positive integer greater than or equal to 3; the element information is related to the tag information and the attribute information; the first triplet information is used to construct a knowledge graph for guiding the generation of marketing text.
[0038] As an optional implementation, in a second aspect of the present invention, the first processing module includes a first processing submodule, a second processing submodule, and a labeling submodule, wherein:
[0039] The first processing submodule is used to clean the marketing text information to be processed using preset text cleaning rules to obtain text information to be segmented; the cleaning process is used to filter non-standard information.
[0040] The second processing submodule is used to perform word segmentation on the text information to be segmented to obtain text information to be classified; the word segmentation is used to reduce the uncertainty of the text information to be segmented.
[0041] The annotation submodule is used to annotate the text information to be classified using a preset classification model to obtain tag information; the tag information includes industry information and text style information.
[0042] As an optional implementation, in a second aspect of the present invention, the second processing module uses preset attribute processing rules to identify and process the marketing text information to be processed, and obtains the attribute information in the following specific way:
[0043] The marketing text information to be processed is processed using a preset recognition model to obtain text output information;
[0044] Determine whether the text output information meets the verification conditions, and obtain the judgment result;
[0045] When the judgment result is yes, the text output information is determined to be attribute information;
[0046] When the judgment result is negative, the marketing text information to be processed is processed using preset new word extraction rules to obtain the attribute information.
[0047] As an optional implementation, in a second aspect of the present invention, the second processing module processes the marketing text information to be processed using preset new word extraction rules to obtain the attribute information in the following specific way:
[0048] The marketing text information to be processed is processed by using a preset new word extraction model to obtain word fragment information;
[0049] The word fragment information is classified and filtered to obtain the attribute information.
[0050] As an optional implementation, in a second aspect of the present invention, the classification model includes a first classification model, and / or a second classification model, and / or a third classification model;
[0051] The annotation submodule uses a preset classification model to annotate the text information to be classified, and the specific method for obtaining the tag information is as follows:
[0052] The industry information is obtained by labeling the text information to be classified using the first classification model;
[0053] The text information to be classified is labeled using the second classification model to obtain the text style information; or,
[0054] The third classification model is used to label the text information to be classified, thereby obtaining the industry information and the text style information.
[0055] As an optional implementation, in a second aspect of the present invention, after the second processing module identifies and processes the marketing text information to be processed using preset attribute processing rules to obtain attribute information, the device further includes:
[0056] The fourth processing module is used to perform word extraction and filtering on the marketing text information to be processed, so as to obtain the word segmentation information to be aggregated.
[0057] The clustering processing module is used to perform clustering processing on the word segmentation information to be clustered to obtain cluster information; the cluster information includes cluster information and category name information.
[0058] The fifth processing module is used to process the cluster information and the attribute information to obtain the second triplet information.
[0059] As an optional implementation, in the second aspect of the present invention, the clustering processing module performs clustering processing on the word segmentation information to be clustered to obtain cluster information in the following specific way:
[0060] The word segmentation information to be clustered is processed using a preset clustering model to obtain the cluster information; the cluster information includes L clusters; where L is a positive integer greater than or equal to 3;
[0061] The cluster information is labeled to obtain the category name information; the category name information includes the L category names.
[0062] A third aspect of the present invention discloses another data processing apparatus based on marketing text, the apparatus comprising:
[0063] Memory containing executable program code;
[0064] A processor coupled to the memory;
[0065] The processor calls the executable program code stored in the memory to execute some or all of the steps in the data processing method based on marketing text disclosed in the first aspect of the present invention.
[0066] The fourth aspect of the present invention discloses a computer storage medium storing computer instructions, which, when invoked, are used to execute some or all of the steps in the data processing method based on marketing text disclosed in the first aspect of the present invention.
[0067] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0068] In this embodiment of the invention, marketing text information to be processed is obtained; the marketing text information to be processed is classified using preset tag processing rules to obtain tag information; the tag information is related to the industry and text style of the marketing text information to be processed; the marketing text information to be processed is identified using preset attribute processing rules to obtain attribute information; the tag information and attribute information are processed to obtain first triplet information; the first triplet information includes M element information; M is a positive integer greater than or equal to 3; the element information is related to the tag information and attribute information; the first triplet information is used to construct a knowledge graph for guiding the generation of marketing text. It can be seen that this invention can obtain tag information and attribute information by processing the marketing text information to be processed through tag processing rules and attribute processing rules, and then comprehensively process them to obtain the first triplet information used to construct a knowledge graph for guiding the generation of marketing text. This is beneficial to improving the tagging efficiency of marketing text, thereby improving the quality of key information extraction, and thus constructing an effective knowledge graph. Attached Figure Description
[0069] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0070] Figure 1 This is a flowchart illustrating a data processing method based on marketing text disclosed in an embodiment of the present invention;
[0071] Figure 2 This is a flowchart illustrating another data processing method based on marketing text disclosed in an embodiment of the present invention;
[0072] Figure 3 This is a schematic diagram of the structure of a data processing device based on marketing text disclosed in an embodiment of the present invention;
[0073] Figure 4 This is a schematic diagram of another data processing device based on marketing text disclosed in an embodiment of the present invention;
[0074] Figure 5 A schematic diagram of the structure of another data processing device based on marketing text disclosed in this embodiment of the invention. Detailed Implementation
[0075] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0076] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.
[0077] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0078] This invention discloses a data processing method and apparatus based on marketing text. It can obtain tag information and attribute information from the marketing text information to be processed through tag processing rules and attribute processing rules, and then comprehensively process them to obtain the first triplet information for constructing a knowledge graph that guides the generation of marketing text. This improves the efficiency of tagging processing of marketing text, thereby enhancing the quality of key information extraction and constructing an effective knowledge graph. Detailed descriptions follow.
[0079] Example 1
[0080] Please see Figure 1 , Figure 1 This is a flowchart illustrating a data processing method based on marketing text disclosed in an embodiment of the present invention. Figure 1 The described data processing method based on marketing text is applied to a data processing system, such as a local server or cloud server for managing marketing text-based data processing; however, this embodiment of the invention is not limited to such applications. Figure 1 As shown, this data processing method based on marketing text may include the following operations:
[0081] 101. Obtain marketing text information to be processed.
[0082] 102. Classify the marketing text information to be processed using the preset tag processing rules to obtain tag information.
[0083] In this embodiment of the invention, the aforementioned tag information is related to the industry and text style of the marketing text information to be processed.
[0084] 103. Use preset attribute processing rules to identify and process the marketing text information to be processed, and obtain attribute information.
[0085] 104. Process the tag information and attribute information to obtain the first triplet information.
[0086] In this embodiment of the invention, the first triplet information includes M element information.
[0087] In this embodiment of the invention, M is a positive integer greater than or equal to 3.
[0088] In this embodiment of the invention, the above-mentioned element information is related to tag information and attribute information.
[0089] In this embodiment of the invention, the aforementioned first triplet information is used to construct a knowledge graph that guides the generation of marketing text.
[0090] Optionally, the above-mentioned label information includes industry information and / or text style information, as described in this embodiment of the invention.
[0091] Optionally, the above industry information includes games, and / or finance, and / or culture and entertainment, and / or e-commerce, and / or skincare and beauty, and / or apparel and footwear, and / or education and training, and / or healthcare, and / or food and beverage, and / or digital home appliances, and / or home environment, and / or social networking and dating, and / or automobiles. This embodiment of the invention does not limit the scope of the invention.
[0092] Optionally, the above text style information includes "finally happened," and / or "must read / play," and / or "free trial," and / or "highlighting offers," and / or "raising questions," and / or "creating scarcity," and / or "encouraging trying," and / or "herd mentality," which are not limited in the embodiments of the present invention.
[0093] Optionally, the above attribute information includes brand information, and / or product attribute information, and / or category information, and / or ingredient information, and / or efficacy information, and / or selling point information. This embodiment of the invention does not limit the scope of the information.
[0094] Optionally, the above first triplet information can be in the following form:
[0095] {First element information, second element information, third element information}
[0096] Preferably, the first element information and the second element information are obtained by filtering from the above-mentioned tag information and attribute information.
[0097] The optional third element information described above represents the relationship between the first element information and the second element information.
[0098] As can be seen, the data processing method based on marketing text described in the embodiments of the present invention can obtain tag information and attribute information by processing the marketing text information to be processed through tag processing rules and attribute processing rules, and then perform comprehensive processing to obtain the first triplet information for constructing a knowledge graph that guides the generation of marketing text. This is beneficial to improving the tagging efficiency of marketing text, thereby improving the quality of key information extraction, so as to construct an effective knowledge graph.
[0099] In an optional embodiment, step 102 above uses preset tagging rules to classify the marketing text information to be processed, obtaining tagging information, including:
[0100] The marketing text information to be processed is cleaned using preset text cleaning rules to obtain the text information to be segmented; the cleaning process is used to filter out non-standard information.
[0101] The text information to be segmented is processed into words to obtain text information to be classified; word segmentation is used to reduce the uncertainty of the text information to be segmented.
[0102] The text information to be classified is labeled using a pre-defined classification model to obtain tag information; the tag information includes industry information and text style information.
[0103] Optionally, the above cleaning process includes removing short text from the marketing text information to be processed, and / or removing non-Chinese characters from the marketing text information to be processed, and / or removing special characters from the marketing text information to be processed. This embodiment of the invention does not limit the scope of the process.
[0104] Optionally, the short text mentioned above includes sentences with a length less than N. Further, N is a positive integer greater than or equal to 2.
[0105] Optionally, the above-mentioned text information to be segmented includes Chinese text information, and / or English text information, and / or numeric text information, which is not limited in the embodiments of the present invention.
[0106] In this optional embodiment, as an optional implementation method, the specific way to perform word segmentation processing on the text information to be segmented to obtain the text information to be classified is as follows:
[0107] The text information to be segmented is imported into the LAC word segmenter for processing to obtain the segmented text information;
[0108] The segmented text information is transformed into numerical vectors to obtain the text information to be classified.
[0109] As can be seen, the data processing method based on marketing text described in the embodiments of the present invention can comprehensively process the marketing text information to be processed through text cleaning rules and classification models to obtain tag information, which is conducive to improving the tagging efficiency of marketing text, thereby improving the quality of key information extraction, so as to construct an effective knowledge graph.
[0110] In another optional embodiment, the above-mentioned identification and processing of the marketing text information to be processed using preset attribute processing rules to obtain attribute information includes:
[0111] The marketing text information to be processed is processed using a preset recognition model to obtain the text output information;
[0112] Determine whether the text output information meets the validation conditions and obtain the result.
[0113] When the judgment result is yes, the text output information is determined to be attribute information;
[0114] When the judgment result is negative, the marketing text information to be processed is processed using the preset new word extraction rules to obtain attribute information.
[0115] Optionally, the above-mentioned method for obtaining the recognition dataset used to train the recognition model is as follows:
[0116] Establish an entity recognition database;
[0117] The training text information is matched using an entity recognition database to obtain matching word information;
[0118] The matched word information is labeled using BIO to obtain BIO word segmentation information;
[0119] The BIO segmentation information is transformed to obtain the first training segmentation information;
[0120] The first training word segmentation information and entity recognition library are processed to obtain the recognition dataset.
[0121] Optionally, the above verification condition is whether the text output information contains BIO format information.
[0122] Optionally, the above judgment result is yes when the text output information contains BIO format information.
[0123] Optionally, any character in the marketing text message can uniquely correspond to B, or I, or O.
[0124] Optionally, B above indicates the beginning of a text information entity segment.
[0125] Optionally, the I above represents the intermediate text information entity.
[0126] Optionally, the O above indicates the end of a text information entity segment.
[0127] As can be seen, the data processing method based on marketing text described in the embodiments of the present invention can comprehensively process the marketing text information to be processed using the recognition model and new word extraction rules to obtain attribute information, which is conducive to improving the tagging efficiency of marketing text and thus improving the quality of key information extraction, so as to construct an effective knowledge graph.
[0128] In another optional embodiment, the marketing text information to be processed is processed using preset new word extraction rules to obtain attribute information, including:
[0129] A pre-defined new word extraction model is used to extract words from the marketing text information to be processed, resulting in word fragment information.
[0130] The word fragment information is classified and filtered to obtain attribute information.
[0131] In this optional embodiment, as an optional implementation method, the specific way to extract words from the marketing text information to be processed using the preset new word extraction model to obtain word fragment information is as follows:
[0132] Break the strings that are not Chinese, English, or numbers in the marketing text information to be processed to obtain character data information;
[0133] The character data information is combined into two characters to obtain a set of two-character information; the set of two-character information includes several two-character information.
[0134] Calculate the mutual information entropy of all two-character information in the two-character information set to obtain a set of mutual information entropy values; the set of mutual information entropy values includes several mutual information entropy values.
[0135] For any mutual information entropy value, determine whether the mutual information entropy value is less than the entropy threshold to obtain the entropy judgment result;
[0136] When the above entropy judgment result is yes, disconnect the two-word information corresponding to the mutual information entropy value to obtain disconnected two-word information, and remove the two-word information corresponding to the mutual information entropy value from the two-word information set to generate a new two-word information set; disconnected two-word information includes several disconnected two-words; any disconnected two-word includes two characters;
[0137] Frequency statistics are performed on all characters in the two-word messages to obtain a character frequency set; the above character frequency set includes several character frequencies;
[0138] For any character frequency, determine whether the character frequency is less than the frequency threshold to obtain the frequency determination result;
[0139] However, if the frequency determination result is yes, the two broken characters corresponding to the frequency of that character will be removed from the information of broken characters, and new information of broken characters will be generated.
[0140] Based on the information of the two separate characters and the set of information of the two characters, the information of the word segments is determined.
[0141] In this optional embodiment, as another optional implementation, the specific method for classifying and filtering the word fragment information to obtain attribute information is as follows:
[0142] The word fragment information is imported into the word classification model for classification to obtain the word information to be filtered;
[0143] The entity recognition library is used to filter the information of the words to be filtered, and the attribute information is obtained.
[0144] As can be seen, the data processing method based on marketing text described in the embodiments of the present invention can use a new word extraction model to extract words from the marketing text information to be processed to obtain word fragment information, and then obtain attribute information through classification and filtering of the word fragment information. This is more conducive to improving the tagging efficiency of marketing text, thereby improving the quality of key information extraction, so as to construct an effective knowledge graph.
[0145] In yet another optional embodiment, the above classification model includes a first classification model, and / or a second classification model, and / or a third classification model;
[0146] The above method uses a pre-defined classification model to label the text information to be classified, obtaining tag information, including:
[0147] The first classification model is used to label the text information to be classified, thereby obtaining industry information;
[0148] The text information to be classified is labeled using a second classification model to obtain text style information; or,
[0149] The third classification model is used to label the text information to be classified, thereby obtaining industry information and text style information.
[0150] Optionally, the above classification models include CNN-based models, and / or FASTTEXT-based models, and / or GRU-based models, and / or LSTM-based models, and / or SVM-based models, and / or BART-based models, and / or BERT-based models, and / or TextCNN-based models. This embodiment of the invention does not limit the specific models.
[0151] Optionally, the third classification model above outputs a combined information, which includes industry information and text style information.
[0152] As can be seen, the data processing method based on marketing text described in the embodiments of the present invention can use a classification model to label the text information to be classified to obtain industry information and text style information, which is more conducive to improving the efficiency of tagging processing of marketing text, thereby improving the quality of key information extraction, so as to construct an effective knowledge graph.
[0153] Example 2
[0154] Please see Figure 2 , Figure 2 This is a flowchart illustrating another data processing method based on marketing text disclosed in an embodiment of the present invention. Figure 2 The described data processing method based on marketing text is applied to a data processing system, such as a local server or cloud server for managing marketing text-based data processing; however, this embodiment of the invention is not limited to such applications. Figure 2 As shown, this data processing method based on marketing text may include the following operations:
[0155] 201. Obtain marketing text information to be processed.
[0156] 202. Classify the marketing text information to be processed using the preset tag processing rules to obtain tag information.
[0157] 203. Use the preset attribute processing rules to identify and process the marketing text information to be processed, and obtain the attribute information.
[0158] 204. Perform word extraction and filtering on the marketing text information to be processed to obtain the word segmentation information to be aggregated.
[0159] 205. Perform clustering processing on the information to be clustered to obtain cluster information.
[0160] In this embodiment of the invention, the cluster information includes cluster information and category name information.
[0161] 206. Process the cluster information and attribute information to obtain the second triplet information.
[0162] 207. Process the tag information and attribute information to obtain the first triplet information.
[0163] In this embodiment of the invention, for the specific technical details and explanations of technical terms for steps 201-203 and 207, please refer to the detailed description of steps 101-104 in Embodiment 1. This embodiment of the invention will not repeat them here.
[0164] In this embodiment of the invention, step 207 can be executed before step 204, after step 204, or in parallel with step 204. This embodiment of the invention does not impose any limitations.
[0165] In this optional embodiment, as an optional implementation method, the specific way of the above-mentioned word prompting filtering process is as follows:
[0166] The marketing text information to be processed is extracted to obtain the first keyword information;
[0167] The first word information is segmented to obtain the second word information;
[0168] The second word information is filtered using the aforementioned attribute information and a preset stop word library to obtain clustered word information.
[0169] Optionally, the above second triplet information can be in the following form:
[0170] {First entity information, second entity information, third entity information}
[0171] Optionally, the first entity information mentioned above is filtered from the cluster information.
[0172] Preferably, the second entity information is selected from the attribute information based on the first entity information.
[0173] The optional third entity information mentioned above is filtered from the category name information based on the first entity information.
[0174] As can be seen, the data processing method based on marketing text described in the embodiments of the present invention can obtain tag information and attribute information by processing the marketing text information to be processed through tag processing rules and attribute processing rules, then perform word extraction and filtering processing on the marketing text information to be processed to obtain word segmentation information to be clustered, and perform clustering processing on the word segmentation information to be clustered to obtain cluster information. Then, by processing the cluster information and attribute information, the second triplet information is obtained, which is conducive to improving the tagging processing efficiency of marketing text, thereby improving the quality of key information extraction, so as to construct an effective knowledge graph.
[0175] In an optional embodiment, the above-described clustering process for the word segmentation information to be clustered to obtain cluster information includes:
[0176] The word segmentation information to be clustered is processed using a pre-defined clustering model to obtain cluster information; the cluster information includes L clusters; L is a positive integer greater than or equal to 3;
[0177] The cluster information is labeled to obtain the category name information; the category name information includes L category names.
[0178] Optionally, the above clustering model is based on the Mini Batch K-Means algorithm.
[0179] Optionally, L is set based on prior information.
[0180] Optionally, the category name information above represents the commonalities of the cluster information. For example, the cluster information includes keywords such as constant temperature, bidirectional fast charging, built-in cable, multi-port output, battery clip, tilt power-off, cartoon / IP, head-shaking function, magnetic power bank, and plug function. The category name information represents the unique features, that is, the unique features are the representation of the commonalities of all the keywords in the above cluster information.
[0181] As can be seen, the data processing method based on marketing text described in the embodiments of the present invention can use a clustering model to process the information to be clustered and segmented to obtain cluster information, and then label it to obtain category name information, which is more conducive to improving the efficiency of marketing text tagging processing, thereby improving the quality of key information extraction, so as to construct an effective knowledge graph.
[0182] Example 3
[0183] Please see Figure 3 , Figure 3 This is a schematic diagram of the structure of a data processing device based on marketing text disclosed in an embodiment of the present invention. Figure 3 The described apparatus can be applied to data processing systems, such as local servers or cloud servers for data processing and management based on marketing texts, and the embodiments of the present invention are not limited thereto. Figure 3 As shown, the device may include:
[0184] Module 301 is used to acquire marketing text information to be processed;
[0185] The first processing module 302 is used to classify the marketing text information to be processed according to preset tag processing rules to obtain tag information; the tag information is related to the industry and text style of the marketing text information to be processed.
[0186] The second processing module 303 is used to identify and process the marketing text information to be processed using preset attribute processing rules to obtain attribute information;
[0187] The third processing module 304 is used to process the tag information and attribute information to obtain the first triplet information; the first triplet information includes M element information; M is a positive integer greater than or equal to 3; the element information is related to the tag information and attribute information; the first triplet information is used to construct a knowledge graph that guides the generation of marketing text.
[0188] It is evident that implementation Figure 3 The described marketing text-based data processing device can obtain tag information and attribute information by processing the marketing text information to be processed through tag processing rules and attribute processing rules, and then perform comprehensive processing to obtain the first triplet information for constructing a knowledge graph that guides the generation of marketing text. This is beneficial to improving the tagging efficiency of marketing text, thereby improving the quality of key information extraction, so as to construct an effective knowledge graph.
[0189] In another alternative embodiment, such as Figure 4 As shown, the first processing module 302 includes a first processing submodule 3021, a second processing submodule 3022, and an annotation submodule 3023, wherein:
[0190] The first processing submodule 3021 is used to clean the marketing text information to be processed using preset text cleaning rules to obtain the text information to be segmented; the cleaning process is used to filter non-standard information.
[0191] The second processing submodule 3022 is used to perform word segmentation on the text information to be segmented to obtain text information to be classified; word segmentation is used to reduce the uncertainty of the text information to be segmented.
[0192] The annotation submodule 3023 is used to annotate the text information to be classified using a preset classification model to obtain label information; the label information includes industry information and text style information.
[0193] It is evident that implementation Figure 4The described marketing text-based data processing device can comprehensively process marketing text information to obtain tag information through text cleaning rules and classification models, which is conducive to improving the tagging efficiency of marketing text and thus improving the quality of key information extraction, so as to construct an effective knowledge graph.
[0194] In yet another alternative embodiment, such as Figure 4 As shown, the second processing module 303 uses preset attribute processing rules to identify and process the marketing text information to be processed, and the specific method for obtaining attribute information is as follows:
[0195] The marketing text information to be processed is processed using a preset recognition model to obtain the text output information;
[0196] Determine whether the text output information meets the validation conditions and obtain the result.
[0197] When the judgment result is yes, the text output information is determined to be attribute information;
[0198] When the judgment result is negative, the marketing text information to be processed is processed using the preset new word extraction rules to obtain attribute information.
[0199] It is evident that implementation Figure 4 The described marketing text-based data processing device can comprehensively process the marketing text information to be processed using recognition models and new word extraction rules to obtain attribute information, which is conducive to improving the efficiency of marketing text tagging and thus improving the quality of key information extraction, so as to construct an effective knowledge graph.
[0200] In yet another alternative embodiment, such as Figure 4 As shown, the second processing module 303 processes the marketing text information to be processed using preset new word extraction rules, and the specific method for obtaining attribute information is as follows:
[0201] A pre-defined new word extraction model is used to extract words from the marketing text information to be processed, resulting in word fragment information.
[0202] The word fragment information is classified and filtered to obtain attribute information.
[0203] It is evident that implementation Figure 4 The described marketing text-based data processing device can use a new word extraction model to extract words from the marketing text information to obtain word fragment information, and then obtain attribute information through classification and filtering of the word fragment information. This is more conducive to improving the tagging efficiency of marketing text, thereby improving the quality of key information extraction and constructing an effective knowledge graph.
[0204] In yet another alternative embodiment, such as Figure 4 As shown, the classification model includes a first classification model, and / or a second classification model, and / or a third classification model;
[0205] The annotation submodule 3023 uses a preset classification model to annotate the text information to be classified, and the specific method for obtaining the label information is as follows:
[0206] The first classification model is used to label the text information to be classified, thereby obtaining industry information;
[0207] The text information to be classified is labeled using a second classification model to obtain text style information; or,
[0208] The third classification model is used to label the text information to be classified, thereby obtaining industry information and text style information.
[0209] It is evident that implementation Figure 4 The described marketing text-based data processing device can use a classification model to label the text information to be classified to obtain industry information and text style information, which is more conducive to improving the efficiency of marketing text labeling and processing, thereby improving the quality of key information extraction and constructing an effective knowledge graph.
[0210] In yet another alternative embodiment, such as Figure 4 As shown, after the second processing module 303 uses preset attribute processing rules to identify and process the marketing text information to be processed and obtains the attribute information, the device further includes:
[0211] The fourth processing module 305 is used to perform word extraction and filtering on the marketing text information to be processed, and obtain the word segmentation information to be aggregated.
[0212] Clustering processing module 306 is used to perform clustering processing on the word segmentation information to be clustered to obtain cluster information; the cluster information includes cluster information and category name information.
[0213] The fifth processing module 307 is used to process the cluster information and attribute information to obtain the second triplet information.
[0214] It is evident that implementation Figure 4 The described marketing text-based data processing device can obtain tag information and attribute information by processing the marketing text information to be processed through tag processing rules and attribute processing rules. Then, it performs word extraction and filtering processing on the marketing text information to be processed to obtain word segmentation information to be clustered. The word segmentation information to be clustered is then clustered to obtain cluster information. Finally, by processing the cluster information and attribute information, the second triplet information is obtained. This is beneficial to improving the tagging efficiency of marketing text, thereby improving the quality of key information extraction and constructing an effective knowledge graph.
[0215] In yet another alternative embodiment, such as Figure 4 As shown, the clustering processing module 306 performs clustering processing on the word segmentation information to be clustered, and obtains the cluster information in the following specific way:
[0216] The word segmentation information to be clustered is processed using a pre-defined clustering model to obtain cluster information; the cluster information includes L clusters; L is a positive integer greater than or equal to 3;
[0217] The cluster information is labeled to obtain the category name information; the category name information includes L category names.
[0218] It is evident that implementation Figure 4 The described marketing text-based data processing device can use a clustering model to process the segmented word information to obtain cluster information, and then label it to obtain category name information. This is more conducive to improving the efficiency of marketing text tagging and processing, thereby improving the quality of key information extraction and constructing an effective knowledge graph.
[0219] Example 4
[0220] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of another data processing device based on marketing text disclosed in an embodiment of the present invention. Figure 5 The described apparatus can be applied to data processing systems, such as local servers or cloud servers for data processing and management based on marketing texts, and the embodiments of the present invention are not limited thereto. Figure 5 As shown, the device may include:
[0221] Memory 401 storing executable program code;
[0222] Processor 402 coupled to memory 401;
[0223] The processor 402 calls the executable program code stored in the memory 401 to perform the steps in the data processing method based on marketing text described in Embodiment 1 or Embodiment 2.
[0224] Example 5
[0225] This invention discloses a computer read storage medium that stores a computer program for electronic data interchange, wherein the computer program causes a computer to perform the steps in the marketing text-based data processing method described in Embodiment 1 or Embodiment 2.
[0226] Example 6
[0227] This invention discloses a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to perform the steps in the marketing text-based data processing method described in Embodiment 1 or Embodiment 2.
[0228] The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0229] Through the detailed description of the above embodiments, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, including read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-Erasable Programmable Read-Only Memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.
[0230] Finally, it should be noted that the data processing method and apparatus based on marketing text disclosed in the embodiments of the present invention are merely preferred embodiments of the present invention and are only used to illustrate the technical solutions of the present invention, not to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data processing method based on marketing text, characterized by, The method comprises: acquiring marketing text information to be processed; classifying the marketing text information to be processed using preset label processing rules to obtain label information; the label information is related to the industry and text style of the marketing text information to be processed; identifying the marketing text information to be processed using preset attribute processing rules to obtain attribute information; processing the label information and the attribute information to obtain first triple information; the first triple information comprises M element information; M is a positive integer greater than or equal to 3; the element information is related to the label information and the attribute information; the first triple information is used to construct a knowledge graph for guiding the generation of marketing text; and the classification of the marketing text information to be processed using preset label processing rules to obtain label information comprises: cleaning the marketing text information to be processed using preset text cleaning rules to obtain text information to be segmented; the cleaning is used to filter non-standard information; segmenting the text information to be segmented to obtain text information to be classified; the segmentation is used to reduce the uncertainty of the text information to be segmented; annotating the text information to be classified using a preset classification model to obtain label information; the label information comprises industry information and text style information; and the identification of the marketing text information to be processed using preset attribute processing rules to obtain attribute information comprises: processing the marketing text information to be processed using a preset identification model to obtain text output information; determining whether the text output information meets a verification condition to obtain a determination result; when the determination result is yes, determining that the text output information is attribute information; when the determination result is no, processing the marketing text information to be processed using preset new word extraction rules to obtain the attribute information; and the processing of the marketing text information to be processed using preset new word extraction rules to obtain the attribute information comprises: extracting words from the marketing text information to be processed using a preset new word extraction model to obtain word piece information; classifying and filtering the word piece information to obtain the attribute information; and the extraction of words from the marketing text information to be processed using a preset new word extraction model to obtain word piece information comprises: breaking a string of non-English and non-Chinese numbers in the marketing text information to be processed to obtain character data information; combining two characters to obtain a two-character information set; the two-character information set comprises a plurality of two-character information; calculating the mutual information entropy of all two-character information in the two-character information set to obtain a mutual information entropy value set; the mutual information entropy value set comprises a plurality of mutual information entropy values; for any mutual information entropy value, determining whether the mutual information entropy value is less than an entropy threshold to obtain an entropy determination result; when the entropy judgment result is yes, disconnecting two-character information corresponding to the mutual information entropy value, obtaining disconnected two-character information, and removing the two-character information corresponding to the mutual information entropy value from the two-character information set to generate a new two-character information set; the disconnected two-character information includes a plurality of disconnected two-character information; any disconnected two-character information includes two characters; frequency statistics are performed on the characters in all the disconnected two-character information to obtain a character frequency set; the character frequency set includes a plurality of character frequencies; for any character frequency, it is judged whether the character frequency is less than a frequency threshold to obtain a frequency judgment result; when the frequency judgment result is yes, the disconnected two-character corresponding to the character frequency is removed from the disconnected two-character information, and new disconnected two-character information is generated; the word piece information is determined according to the disconnected two-character information and the two-character information set.
2. The data processing method based on marketing text according to claim 1, characterized in that, The classification model includes a first classification model, and / or a second classification model, and / or a third classification model; The method further includes: The industry information is obtained by using the first classification model to label the text information to be classified; The text style information is obtained by using the second classification model to label the text information to be classified; or The industry information and the text style information are obtained by using the third classification model to label the text information to be classified. 3.The marketing text-based data processing method of claim 1, wherein, After the attribute information is obtained by using the preset attribute processing rule to identify and process the text information to be processed, the method further includes: The text information to be processed is filtered to obtain text information to be clustered; The text information to be clustered is clustered to obtain cluster information; the cluster information includes cluster information and category name information; The cluster information and the attribute information are processed to obtain second triple information.
4. The data processing method based on marketing text according to claim 3, characterized in that, The method further includes: The cluster information is obtained by using a preset clustering model to process the text information to be clustered; the cluster information includes L clusters; L is a positive integer greater than or equal to 3; The category name information is obtained by labeling the cluster information; the category name information includes L category names.
5. A data processing apparatus based on marketing text, characterized by, The device includes: An acquisition module is configured to acquire text information to be processed; A first processing module is configured to use a preset label processing rule to classify the text information to be processed to obtain label information; the label information is related to the industry and the text style of the text information to be processed; A second processing module is configured to use a preset attribute processing rule to identify and process the text information to be processed to obtain attribute information. The third processing module is configured to process the label information and the attribute information to obtain first triple information; the first triple information comprises M element information; M is a positive integer greater than or equal to 3; the element information is related to the label information and the attribute information; and the first triple information is used to construct a knowledge graph for guiding generation of a marketing text. The first processing module comprises a first processing submodule, a second processing submodule and a labeling submodule. The first processing submodule is configured to clean the to-be-processed marketing text information by using a preset text cleaning rule to obtain to-be-processed text information for word segmentation; and the cleaning is used to filter non-standard information. The second processing submodule is configured to perform word segmentation processing on the to-be-processed text information to obtain to-be-classified text information; and the word segmentation processing is used to reduce the uncertainty of the to-be-processed text information. The labeling submodule is configured to label the to-be-classified text information by using a preset classification model to obtain label information; and the label information comprises industry information and text style information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The second processing module is configured to identify the to-be-processed marketing text information by using a preset attribute processing rule to obtain attribute information. The Counting frequencies of all characters in the disconnection two-character information to obtain a character frequency set; the character frequency set includes a plurality of character frequencies; For any character frequency, determining whether the character frequency is less than a frequency threshold to obtain a frequency determination result; When the frequency determination result is yes, removing the disconnection two-character corresponding to the character frequency from the disconnection two-character information, and generating new disconnection two-character information; Determining word piece information according to the disconnection two-character information and the two-character information set.
6. A data processing apparatus based on marketing text, characterized by, The apparatus includes: a memory storing executable program codes; a processor coupled with the memory; the processor invokes the executable program codes stored in the memory to execute the data processing method based on marketing text according to any one of claims 1-4.
7. A computer storable medium, characterized by The computer storage medium stores computer instructions, which are invoked to execute the data processing method based on marketing text according to any one of claims 1-4.
Citation Information
Patent Citations
Character search system and method based on search engines
CN107908749A
Context recognition complementing method and system based on knowledge graph, terminal and medium
CN109657238A
Feature word processing method and device in content delivery system and storage medium
CN110020120A
Intelligent customer service processing method and system, and equipment
CN111782793A