Text processing method, electronic device and readable storage medium
By extracting the target opinion phrases in the comment text and performing template replacement processing, the problem of users' difficulty in obtaining comment information intuitively is solved, and the application effect of opinion phrases and user reading experience is improved.
Patent Information
- Application Number
- CN202210453232.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-24
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-04-24
AI Technical Summary
When users read comment texts, it is difficult for users to intuitively obtain effective comment information about entities, resulting in poor application of opinion phrases and poor user reading experience.
By extracting the target view phrases in the comment text, and based on the preset view word template, selecting the target template corresponding to the target entity word, and replacing the target placeholder with the target view word to form the processing result.
The normalization of opinion phrases in the comment text is realized, which improves the user's reading experience of opinion phrases, allowing users to obtain effective comment information about entities more intuitively.
Smart Images

Figure CN114780684B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a text processing method, an electronic device and a readable storage medium. Background Art
[0002] At present, with the rapid development of the Internet, people's lives are becoming more and more comfortable with the help of the Internet. People can communicate and shop through the Internet. In the process of using the Internet, users can post comments to express their opinions, and users can also obtain other users' opinions on a certain thing by reading comments. For example, when shopping through the Internet, since merchants and users cannot communicate face to face, merchants can understand user needs and other information through the comments posted by users, and improve the products based on the comments posted by users to make the products better meet user needs.
[0003] In the related art, the method of processing the comment text and obtaining the user's opinion is usually to extract the opinion phrases in the comment text through rule matching or sequence labeling, and obtain the user's opinion based on the extracted opinion phrases. The opinion phrases include entity words and description texts describing the entity corresponding to the entity words. The so-called entity words usually include nouns and pronouns, such as product names, product brands, etc.
[0004] However, in the related art, since the comment text posted by the user is relatively complex, the opinion phrases obtained for the same entity are also relatively complex, and the user cannot intuitively obtain effective comment information about the entity based on the opinion phrases. As a result, the application effect of the opinion phrases is poor, and thus the user's reading experience of the extracted opinion phrases is poor. Summary of the invention
[0005] The purpose of the embodiments of the present invention is to provide a text processing method, an electronic device and a readable storage medium to improve the user's reading experience of the extracted opinion phrases. The specific technical solution is as follows:
[0006] In a first aspect, an embodiment of the present invention provides a text processing method, the method comprising:
[0007] Get the comment text to be processed;
[0008] Extracting a target opinion phrase from the review text to be processed, and based on the positional relationship between the target entity word in the target opinion phrase and the target description text, selecting a target template corresponding to the target entity word in the target opinion phrase from a preset opinion word template; wherein the target description text is used to describe the entity corresponding to the target entity word, each opinion word template includes: an entity word and a placeholder, and the placeholder is associated with at least one candidate opinion word, and the candidate opinion word included in each opinion word template is used to describe the entity corresponding to the entity word included in the opinion word template;
[0009] Based on the correspondence between the description text and the opinion word preset for each entity word, a target opinion word corresponding to the target description text is selected from at least one candidate opinion word associated with the target placeholder word in the target template;
[0010] The target placeholder is replaced by the target opinion word, and the replaced target template is determined as a processing result of the comment text to be processed.
[0011] Optionally, in a specific implementation, the extracting a target opinion phrase from the to-be-processed comment text, and selecting a target template corresponding to the target entity word in the target opinion phrase from preset opinion word templates based on the positional relationship between the target entity word in the target opinion phrase and the target description text, includes:
[0012] Inputting the review text to be processed into a preset sequence labeling model, and obtaining a target opinion phrase output by the sequence labeling model and a target template corresponding to the target opinion phrase;
[0013] Among them, the sequence labeling model is obtained based on the training of multiple preset first samples; the first sample includes: a first sample opinion phrase marked with a starting position label and an intermediate position label, the starting position label includes: a starting identifier for representing the beginning of the first sample opinion phrase and a sample template identifier of the first sample template corresponding to the first sample opinion phrase; the intermediate position label includes: an intermediate label for representing the end of the first sample opinion phrase and the sample template identifier.
[0014] Optionally, in a specific implementation, based on the correspondence between the description text and the opinion word preset for each entity word, selecting a target opinion word corresponding to the target description text from at least one candidate opinion word associated with the target placeholder word in the target template includes:
[0015] Inputting the target opinion phrase and the target template into a concatenated text into a preset opinion confirmation model, and obtaining the target opinion word output by the opinion confirmation model;
[0016] Wherein, the opinion confirmation model is obtained by training based on a plurality of preset second samples; the second sample includes: a second sample opinion phrase and a second sample template corresponding to the second sample opinion phrase, the second sample template includes: a second sample entity word corresponding to the same entity as the entity word in the second sample opinion phrase, and a sample opinion word selected from at least one candidate opinion word associated with the second sample placeholder in the second sample template, corresponding to the description text in the second sample opinion phrase and used to replace the second sample placeholder.
[0017] Optionally, in a specific implementation, the opinion confirmation model includes: a classification layer; the step of inputting the concatenated text obtained by concatenating the target opinion phrase and the target template into a preset opinion confirmation model, and obtaining the target opinion word output by the opinion confirmation model includes:
[0018] Inputting the concatenated text obtained by concatenating the target opinion phrase and the target template into a preset opinion confirmation model to obtain a latent vector of the target placeholder determined based on the target opinion phrase and the specified content in the target template; wherein the specified content is: the content in the target template other than the target placeholder;
[0019] The latent vector is input to the classification layer so that the classification layer classifies the latent vector based on at least one candidate opinion word associated with the target placeholder in the target template, and outputs the candidate opinion word representing the category to which the latent vector belongs as the target opinion word.
[0020] Optionally, in a specific implementation, the method of constructing the opinion word template includes:
[0021] Obtaining initial opinion phrases that appear in the comment text to be analyzed at a frequency higher than a specified frequency;
[0022] Replacing the opinion word in the initial opinion phrase with a preset placeholder, and determining the replaced opinion word as a candidate opinion word associated with the placeholder, to obtain an initial template;
[0023] The initial templates for the same entity and including entity words and placeholder words with the same positional relationship are merged to obtain various opinion word templates.
[0024] Optionally, in a specific implementation, obtaining the comment text to be processed includes:
[0025] An initial comment text is obtained, and data of the initial comment text is initialized to obtain a comment text to be processed.
[0026] Optionally, in a specific implementation, the step of initializing the data of the initial comment text to obtain the comment text to be processed includes:
[0027] The initial comment text is cleaned and the initial comment text after data cleaning is subjected to designated processing to obtain the comment text to be processed; wherein the designated processing includes: word segmentation processing and / or character segmentation processing.
[0028] In a second aspect, an embodiment of the present invention provides a text processing device, the device comprising:
[0029] A text acquisition module is used to obtain the comment text to be processed;
[0030] A target template determination module is used to extract a target opinion phrase from the comment text to be processed, and based on the positional relationship between the target entity word in the target opinion phrase and the target description text, select a target template corresponding to the target entity word in the target opinion phrase from a preset opinion word template; wherein the target description text is used to describe the entity corresponding to the target entity word, each opinion word template includes: an entity word and a placeholder, and the placeholder is associated with at least one candidate opinion word, and the candidate opinion word included in each opinion word template is used to describe the entity corresponding to the entity word included in the opinion word template;
[0031] a target opinion word determination module, configured to select a target opinion word corresponding to the target description text from at least one candidate opinion word associated with the target placeholder word in the target template based on a correspondence relationship between the description text and the opinion word preset for each entity word;
[0032] A result acquisition module is used to replace the target placeholder word with the target opinion word, and determine the replaced target template as a processing result of the comment text to be processed.
[0033] Optionally, in a specific implementation, the target template determination module is specifically used to:
[0034] Inputting the review text to be processed into a preset sequence labeling model, and obtaining a target opinion phrase output by the sequence labeling model and a target template corresponding to the target opinion phrase;
[0035] Among them, the sequence labeling model is obtained based on the training of multiple preset first samples; the first sample includes: a first sample opinion phrase marked with a starting position label and an intermediate position label, the starting position label includes: a starting identifier for representing the beginning of the first sample opinion phrase and a sample template identifier of the first sample template corresponding to the first sample opinion phrase; the intermediate position label includes: an intermediate label for representing the end of the first sample opinion phrase and the sample template identifier.
[0036] Optionally, in a specific implementation, the target opinion word determination module includes:
[0037] A target opinion word acquisition submodule is used to input the concatenated text obtained by concatenating the target opinion phrase and the target template into a preset opinion confirmation model, and obtain the target opinion word output by the opinion confirmation model;
[0038] Wherein, the opinion confirmation model is obtained by training based on a plurality of preset second samples; the second sample includes: a second sample opinion phrase and a second sample template corresponding to the second sample opinion phrase, the second sample template includes: a second sample entity word corresponding to the same entity as the entity word in the second sample opinion phrase, and a sample opinion word selected from at least one candidate opinion word associated with the second sample placeholder in the second sample template, corresponding to the description text in the second sample opinion phrase and used to replace the second sample placeholder.
[0039] Optionally, in a specific implementation, the opinion confirmation model includes: a classification layer; the target opinion word acquisition submodule is specifically used for:
[0040] Inputting the concatenated text obtained by concatenating the target opinion phrase and the target template into a preset opinion confirmation model, obtaining a latent vector of the placeholder in the target template determined based on the target opinion phrase and the specified content in the target template; wherein the specified content is: the content in the target template other than the target placeholder;
[0041] The latent vector is input to the classification layer so that the classification layer classifies the latent vector based on at least one candidate opinion word associated with the target placeholder in the target template, and outputs the candidate opinion word representing the category to which the latent vector belongs as the target opinion word.
[0042] Optionally, in a specific implementation, the device further includes:
[0043] The template acquisition module is used to obtain initial opinion phrases whose appearance frequency in the comment text to be analyzed is higher than a specified frequency; replace the opinion words in the initial opinion phrases with preset placeholders, and determine the replaced opinion words as candidate opinion words associated with the placeholders to obtain an initial template; merge the initial templates for the same entity and including the same positional relationship between the entity words and the placeholders to obtain individual opinion word templates.
[0044] Optionally, in a specific implementation, the text acquisition module includes:
[0045] The initial text acquisition submodule is used to obtain the initial comment text;
[0046] The initialization submodule is used to perform data initialization on the initial comment text to obtain the comment text to be processed.
[0047] Optionally, in a specific implementation, the initialization submodule is specifically used to:
[0048] The initial comment text is cleaned and the initial comment text after data cleaning is subjected to designated processing to obtain the comment text to be processed; wherein the designated processing includes: word segmentation processing and / or character segmentation processing.
[0049] In a third aspect, an embodiment of the present invention provides an electronic device, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus;
[0050] Memory, used to store computer programs;
[0051] The processor is used to implement the steps of any text processing method provided in the first aspect when executing the program stored in the memory.
[0052] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any text processing method provided in the first aspect are implemented.
[0053] In a fifth aspect, an embodiment of the present invention provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the steps of any text processing method provided in the first aspect.
[0054] Beneficial effects of the embodiments of the present invention:
[0055] As can be seen from the above, when the scheme provided by the embodiment of the present invention is applied, when processing the text, first, the comment text to be processed is obtained, and then the target opinion phrase is extracted from the comment text to be processed, and based on the positional relationship between the target entity word and the target description text in the target opinion phrase, the target template corresponding to the target entity word is selected from the preset opinion word template. Since each opinion word template includes: an entity word and a placeholder, and the placeholder is associated with at least one candidate opinion word for describing the entity corresponding to the entity word included in the opinion word template, the selected target template includes: a target entity word and a target placeholder, and the target placeholder is associated with at least one candidate opinion word. Then, based on the preset correspondence relationship between the description text and the opinion word for each entity word, the target opinion word corresponding to the target description text can be selected from at least one candidate opinion word associated with the target placeholder in the selected target template, so that the target opinion word can be used to replace the target placeholder, and the replaced target template is determined as the processing result of the comment text to be processed.
[0056] Based on this, by applying the solution provided by the embodiment of the present invention, with the help of a preset opinion word template and the preset correspondence between the description text and the opinion word for each entity word, the description text in the target opinion phrase for the same entity extracted from the above-mentioned text to be processed can be unified into the same opinion word, thereby realizing normalization of the description text, making it easier for users to intuitively obtain valid comment information about the entity based on the opinion words and perform subsequent operations on the above-mentioned valid comment information, thereby improving the intuitiveness of the text and improving the user's reading experience of the extracted opinion phrases. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other embodiments can also be obtained based on these drawings.
[0058] Figure 1 A flowchart of a text processing method provided by an embodiment of the present invention;
[0059] Figure 2 It is a structural diagram of a sequence labeling model;
[0060] Figure 3 A schematic diagram of the structure of the model to confirm a point of view;
[0061] Figure 4 A flowchart of another text processing method provided by an embodiment of the present invention;
[0062] Figure 5 A schematic diagram of a specific embodiment of the present invention;
[0063] Figure 6 A structural schematic diagram of a text processing device provided by an embodiment of the present invention;
[0064] Figure 7 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0065] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field based on this application belong to the scope of protection of the present invention.
[0066] In the related art, the method of processing the comment text and obtaining the user's opinion is usually to extract the opinion phrases in the comment text through rule matching or sequence labeling, and obtain the user's opinion based on the extracted opinion phrases. However, in the related art, since the comment text posted by the user is relatively complex, the opinion phrases obtained for the same entity are also relatively complex, and the user cannot intuitively obtain effective comment information about the entity based on the opinion phrases. As a result, the application effect of the opinion phrases is poor, and thus the user's reading experience of the extracted opinion phrases is poor.
[0067] In order to solve the above technical problem, an embodiment of the present invention provides a text processing method.
[0068] Among them, the method can be applied to various application scenarios that require text processing, for example, a merchant needs to process the comment text in order to improve the product according to user needs, and for another example, a user needs to process the comment text in order to read the comment text that he is interested in, etc. In addition, the method can be applied to various electronic devices such as laptops, tablet computers, and mobile phones, hereinafter referred to as electronic devices. Based on this, the embodiment of the present invention does not limit the application scenario and execution subject of the method.
[0069] A text processing method provided by an embodiment of the present invention may include the following steps:
[0070] Get the comment text to be processed;
[0071] A target opinion phrase is extracted from the comment text to be processed, and based on the positional relationship between the target entity word in the target opinion phrase and the target description text, a target template corresponding to the target entity word in the target opinion phrase is selected from preset opinion word templates; wherein the target description text is used to describe the entity corresponding to the target entity word, each opinion word template includes: an entity word and a placeholder, and at least one candidate opinion word associated with the placeholder, and the candidate opinion word included in each opinion word template is used to describe the entity corresponding to the entity word included in the opinion word template;
[0072] Based on the correspondence between the description text and the opinion word preset for each entity word, a target opinion word corresponding to the target description text is selected from at least one candidate opinion word associated with the target placeholder word in the target template;
[0073] The target placeholder is replaced by the target opinion word, and the replaced target template is determined as a processing result of the comment text to be processed.
[0074] As can be seen from the above, when the scheme provided by the embodiment of the present invention is applied, when processing the text, first, the comment text to be processed is obtained, and then the target opinion phrase is extracted from the comment text to be processed, and based on the positional relationship between the target entity word and the target description text in the target opinion phrase, the target template corresponding to the target entity word is selected from the preset opinion word template. Since each opinion word template includes: an entity word and a placeholder, and the placeholder is associated with at least one candidate opinion word for describing the entity corresponding to the entity word included in the opinion word template, the selected target template includes: a target entity word and a target placeholder, and the target placeholder is associated with at least one candidate opinion word. Then, based on the preset correspondence relationship between the description text and the opinion word for each entity word, the target opinion word corresponding to the target description text can be selected from at least one candidate opinion word associated with the target placeholder in the selected target template, so that the target opinion word can be used to replace the target placeholder, and the replaced target template is determined as the processing result of the comment text to be processed.
[0075] Based on this, by applying the solution provided by the embodiment of the present invention, with the help of a preset opinion word template and the preset correspondence between the description text and the opinion word for each entity word, the description text in the target opinion phrase for the same entity extracted from the above-mentioned text to be processed can be unified into the same opinion word, thereby realizing normalization of the description text, making it easier for users to intuitively obtain valid comment information about the entity based on the opinion words and perform subsequent operations on the above-mentioned valid comment information, thereby improving the intuitiveness of the text and improving the user's reading experience of the extracted opinion phrases.
[0076] A text processing method provided by an embodiment of the present invention is described in detail below in conjunction with the accompanying drawings.
[0077] Figure 1 A flowchart of a text processing method provided by an embodiment of the present invention is shown in FIG. Figure 1 As shown, the method may include the following steps S101-S104.
[0078] S101: Obtain the comment text to be processed.
[0079] When processing text, first, obtain the comment text to be processed.
[0080] For example, when processing the comment text of an e-commerce company, it is necessary to obtain the comment text to be processed of the e-commerce company. For another example, when processing the comment text on the Internet, it is necessary to obtain the comment text to be processed on the Internet.
[0081] The initial comment text obtained on the Internet is usually a paragraph or a sentence, and typos, invalid characters, etc. may appear in the initial comment text. Therefore, after obtaining the above initial comment text, the initial comment text can be preprocessed to obtain the comment text to be processed.
[0082] Therefore, optionally, in a specific implementation, the above step S101 may include the following step 11:
[0083] Step 11: Obtain the initial comment text and perform data initialization on the initial comment text to obtain the comment text to be processed.
[0084] In this specific implementation, the initial comment text is first obtained, and then data initialization processing is performed on the obtained initial comment text to obtain the comment text to be processed, so that the obtained text to be processed is more convenient for subsequent processing.
[0085] Optionally, in a specific implementation, the above step 11 may include the following step 21:
[0086] Step 21: Perform data cleaning on the initial comment text, and perform specified processing on the initial comment text after data cleaning to obtain the comment text to be processed.
[0087] The designated processing includes: word segmentation processing and / or character segmentation processing.
[0088] In this specific implementation, the initial comment text is first obtained, then data cleaning is performed on the initial comment text, and word segmentation and / or character segmentation are performed on the initial comment text after data cleaning, thereby obtaining the comment text to be processed.
[0089] Among them, data cleaning refers to the re-examination and verification of data, the purpose of which is to delete duplicate information, correct existing errors, and provide data consistency. For example, if there are repeated words or text errors in an initial comment text, the initial comment text will be data cleaned.
[0090] Word segmentation and / or character segmentation refers to dividing the content included in the above-mentioned initial comment text into the form of words and / or phrases. For example, when processing the initial comment text, the above-mentioned initial comment text is subjected to word segmentation and / or character segmentation to obtain a comment text to be processed that is easy to process.
[0091] For example, the initial comment text is "The night vision picture is extremely blurry", and there are repeated words in the initial text. After data cleaning of the above initial comment text, the comment text obtained is "The night vision picture is extremely blurry". Furthermore, the above "The night vision picture is extremely blurry" is segmented, and the comment text to be processed obtained is "night vision picture", "especially" and "blurred".
[0092] S102: extracting a target opinion phrase from the comment text to be processed, and based on the positional relationship between the target entity word in the target opinion phrase and the target description text, selecting a target template corresponding to the target entity word in the target opinion phrase from preset opinion word templates.
[0093] Among them, the target description text is used to describe the entity corresponding to the target entity word, each opinion word template includes: an entity word and a placeholder word, and the placeholder word is associated with at least one candidate opinion word, and the candidate opinion words included in each opinion word template are used to describe the entity corresponding to the entity word included in the opinion word template.
[0094] When performing text processing on the comment text to be processed, multiple opinion word templates can be pre-constructed; wherein each opinion word template includes: an entity word and a placeholder word, and the placeholder word is associated with at least one candidate opinion word, and the candidate opinion word associated with the placeholder included in each opinion word template is used to describe the entity corresponding to the entity word included in the opinion word template.
[0095] Among them, for the same entity word, multiple different opinion word templates about the entity word can be constructed according to the positional relationship between the entity words and placeholders included. That is to say, for different opinion word templates including the same entity word, the positional relationship between the entity words and placeholders included therein is different.
[0096] For example, the opinion word template "night vision picture [MASK]: clear, blurry"; and the opinion word template "[MASK] night vision picture: clear, blurry" are different opinion word templates including the same entity words; wherein, night vision picture is the entity word, [MASK] is the placeholder word, and clear and blurry is at least one candidate opinion word associated with the placeholder [MASK].
[0097] In this way, after obtaining the comment text to be processed, the target opinion phrase can be extracted from the comment text to be processed. The target opinion phrase includes a target entity word and a target description text for describing the entity corresponding to the target entity word.
[0098] For example, if the target opinion phrase extracted from the comment text to be processed is "the night vision picture is blurry", then the target entity word included in the target opinion phrase is "night vision picture", and the target description text used to describe the entity corresponding to the above target entity word is "blurry". For another example, if the target opinion phrase extracted from the comment text to be processed is "the network connection is often disconnected", then the target entity word included in the target opinion phrase is "network connection", and the target description text used to describe the entity corresponding to the above target entity word is "frequent disconnections".
[0099] Generally, for an entity, the entity can be described based on its functional effects, usage experience, usage characteristics, etc. For example, a computer monitor can be described based on its picture effects, picture size, etc.
[0100] After the target opinion phrase is extracted, the positional relationship between the target entity words and the target description text included in the target opinion phrase may be further determined.
[0101] For example, the target opinion phrase is "the night vision picture is blurry", in which the target entity word is "night vision picture", the target description text is "blurry", and the target entity word "night vision picture" is located before the target description text "blurry"; the target opinion phrase is "blurry night vision picture", in which the target entity word is "night vision picture", the target description text is "blurry", and the target entity word "night vision picture" is located after the target description text "blurry".
[0102] Furthermore, based on the positional relationship between the target entity word in the target opinion phrase and the target description text, a target template corresponding to the target entity word in the target opinion phrase can be selected from the preset opinion word template, and the target template includes the target entity word and the target placeholder word, and the target placeholder word is associated with at least one candidate opinion word.
[0103] Among them, the entity words in the selected target template and the target entity words in the target opinion phrase correspond to the same entity, and the positional relationship between the entity words and placeholders included in the selected target template is the same as the positional relationship between the target entity words and the target description text included in the above-mentioned target opinion phrase, that is, when the target entity word is located before the target description text in the target opinion phrase, the entity word is located before the placeholder in the selected target template; when the target entity word is located after the target description text in the target opinion phrase, the entity word is located after the placeholder in the selected target template.
[0104] S103: Based on the correspondence between the description text and the opinion word preset for each entity word, a target opinion word corresponding to the target description text is selected from at least one candidate opinion word associated with the target placeholder word in the target template.
[0105] After determining the target template corresponding to the target entity word in the above-mentioned target opinion phrase, the target opinion word corresponding to the above-mentioned target description text can be selected from at least one candidate opinion word associated with the target placeholder word in the above-mentioned target template based on the preset correspondence between the description text and the opinion word for each entity word.
[0106] For example, the preset correspondence between the description text and opinion words for "night vision picture" includes: the description text "there are snowflakes" corresponds to the opinion word "blurred", and the description text "the picture is clear" corresponds to the opinion word "clear". Therefore, for the target opinion phrase "there are snowflakes in the night vision picture", after determining that the target entity word "night vision picture" in the target opinion phrase corresponds to the target template, it is possible to select the candidate opinion word "blurred" corresponding to the target description text "there are snowflakes" in the above target opinion phrase from the candidate opinion words associated with the target placeholder words in the target template: "blurred" and "clear" based on the above correspondence, and use the selected "blurred" as the target opinion word corresponding to the above target description text "there are snowflakes".
[0107] For the same entity, due to external factors such as user's language habits and regional dialects, the same opinion word describing the entity may have different expressions. In other words, the same opinion word may correspond to different description texts. For example, for the opinion word "low memory", there may be different description texts such as "small memory", "little memory", "insufficient memory" and so on.
[0108] In this way, in view of the situation that the same opinion word can correspond to different description texts as mentioned above, the correspondence between the description text and the opinion word for each entity word can be pre-set. For example, the description texts for the entity word "night vision picture" include "the screen has snow" and "the screen has stripes", and "the screen has snow" and "the screen has stripes" are two different description texts for the opinion word "blurred". Therefore, for the entity word "night vision picture", the correspondence between the opinion word "blurred" and the description texts "the screen has snow" and "the screen has stripes" can be established.
[0109] Optionally, the above correspondence may be established by collecting a large amount of comment texts and classifying the descriptive texts and opinion words in the collected comment texts, thereby establishing a correspondence between the descriptive texts and opinion words for each entity word based on the classification results.
[0110] Optionally, the above correspondence relationship may be established through manual operation. A person may establish a correspondence relationship between the description text and the opinion word for each entity word based on experience.
[0111] S104: replacing the target placeholder word with the target opinion word, and determining the replaced target template as the processing result of the comment text to be processed.
[0112] After determining the target viewpoint words corresponding to the target description text, the target viewpoint words may be used to replace the target placeholder words in the target template, and the replaced target template may be determined as the processing result of the text to be processed.
[0113] In this way, the processing results of the text to be processed include: the entity words in the target template and the determined target opinion words, and when the target template includes other content besides entity words and placeholder words, the processing results of the text to be processed also include the above-mentioned other content.
[0114] Optionally, in order to make the processing result of the text to be processed a text that conforms to the user's language habits and is fluent, the preset opinion word templates may include, in addition to entity words and placeholders, prepositions, conjunctions, adjectives, etc. to enrich the text and make the text fluent.
[0115] For example, the target template is: The night vision picture of the night vision goggles of model A is very [MASK]: clear, blurry, where night vision picture is the entity word in the target template, [MASK] is the placeholder word in the target template, and clear and blurry are at least one candidate opinion word associated with the placeholder word [MASK]. When the final target opinion word is: clear, the processing result of the text to be processed is: The night vision picture of the night vision goggles of model A is very clear.
[0116] For each entity word that corresponds to a different entity, the opinion word template corresponding to each entity word may be different, and for each entity word that corresponds to the same entity, since the positional relationship between the entity word and the description text used to describe the entity corresponding to the entity word in the opinion phrase where each entity word is located is different, the opinion word template corresponding to each entity word may also be different. Therefore, when constructing an opinion word template, multiple different templates can be constructed for the same entity word.
[0117] Based on this, optionally, in a specific implementation, the method of constructing the opinion word template may include steps 31-33:
[0118] Step 31: obtaining initial opinion phrases in the comment text to be analyzed that appear at a frequency higher than a specified frequency;
[0119] Step 32: replacing the opinion word in the initial opinion phrase with a preset placeholder, and determining the replaced opinion word as a candidate opinion word associated with the placeholder, to obtain an initial template;
[0120] Step 33: merge the initial templates for the same entity and including entity words and placeholder words with the same positional relationship to obtain individual opinion word templates.
[0121] In this specific implementation, in the process of constructing the above-mentioned opinion word template, the initial opinion phrase whose appearance frequency in the comment text to be analyzed is higher than the specified frequency can be first obtained, and then, for the above-mentioned initial opinion phrase, the opinion word in the above-mentioned initial opinion phrase is replaced with a preset placeholder, and the above-mentioned opinion word to be replaced is determined as the candidate opinion word associated with the above-mentioned placeholder, thereby obtaining the initial template, and determining the positional relationship between the entity word and the placeholder in each initial template, and then, the initial templates for the same entity and including the same positional relationship between the entity word and the placeholder are merged to further obtain each opinion word template.
[0122] For example, the initial opinion phrase is "the night vision picture is blurry", and the opinion word in the initial opinion phrase is "blurry". Therefore, the opinion word "blurry" is replaced with a preset placeholder, and the replaced opinion word "blurry" is determined as a candidate opinion word associated with the above placeholder. Thus, the initial template "night vision picture [MASK]: blurry" can be obtained, and in the initial template, the entity word is located before the placeholder.
[0123] For another example, the initial opinion phrase is "The night vision picture is clear", and the opinion word in the initial opinion phrase is "clear". Therefore, the opinion word "clear" is replaced with a preset placeholder, and the replaced opinion word "clear" is determined as a candidate opinion word associated with the above placeholder. Thus, the initial template "Night vision picture [MASK]: clear" can be obtained, and in this initial template, the entity word is located before the placeholder.
[0124] It can be seen from the above that there may be different opinion words for the same entity, which may lead to the same opinion word template for the same entity. Therefore, in order to avoid duplication of opinion word templates, the initial templates for the same entity and with the same positional relationship between the entity words and placeholder words are merged to obtain individual opinion word templates.
[0125] Exemplarily, there are initial templates “night vision screen [MASK]: blurred”, “night vision screen [MASK]: clear”, “[MASK] night vision screen: blurred” and “[MASK] night vision screen: clear”, and the initial templates “night vision screen [MASK]: blurred” and “night vision screen [MASK]: clear” can be merged into the opinion word template “night vision screen [MASK]: blurred, clear”; the initial templates “[MASK] night vision screen: blurred” and “[MASK] night vision screen: clear” can be merged into the opinion word template “[MASK] night vision screen: blurred, clear” to obtain multiple opinion word templates about the same entity word “night vision screen”.
[0126] Optionally, the opinion word template may be constructed by manual operation.
[0127] Optionally, when establishing an opinion word template for a certain vertical field, the opinion word template may be a template for entity words in the vertical field.
[0128] The so-called vertical field refers to providing specific services to a limited group. For example, if you want to establish a viewpoint word template in the field of computers, the viewpoint word template in this field can be: the display is clear or blurry, the CPU runs fast or slow, etc.
[0129] From the above, it can be seen that by applying the solution provided by the embodiment of the present invention, the description texts in the target opinion phrases for the same entity extracted from the above-mentioned text to be processed can be unified into the same opinion word with the help of a preset opinion word template and the correspondence between the opinion words about the description text preset for each entity word, thereby realizing the normalization of the description text, making it easier for users to intuitively obtain valid comment information about the entity based on the opinion words and perform subsequent operations on the above-mentioned valid comment information, thereby improving the intuitiveness of the text and improving the user's reading experience of the extracted opinion phrases.
[0130] Optionally, in a specific implementation, the above step S102: extracting the target opinion phrase from the comment text to be processed, and selecting the target template corresponding to the target entity word in the target opinion phrase from the preset opinion word template based on the positional relationship between the target entity word in the target opinion phrase and the target description text, may include step 41:
[0131] Step 41: Input the comment text to be processed into a preset sequence labeling model, and obtain the target opinion phrase output by the sequence labeling model and the target template corresponding to the target opinion phrase.
[0132] Among them, the sequence labeling model is obtained based on the training of multiple preset first samples; the first sample includes: a first sample opinion phrase annotated with a starting position label and an intermediate position label, the starting position label includes: a starting identifier for representing the beginning of the first sample opinion phrase and a sample template identifier of the first sample template corresponding to the first sample opinion phrase; the intermediate position label includes: an intermediate label for representing the end of the first sample opinion phrase and a sample template identifier.
[0133] In this specific implementation, first, the above-mentioned opinion word template is established, and then the above-mentioned opinion word template can be used to annotate the first sample opinion phrase to obtain the first sample, and then the obtained first sample can be used for model training to obtain a sequence labeling model.
[0134] The labels annotated with each first sample opinion phrase include: a starting position label and a middle position label.
[0135] For example, the comment text to be processed is input into the sequence labeling model. In the sequence labeling model, the comment text to be processed can be labeled by two types of labels "B" and "I", wherein the label "B-template 1" indicates that the current position in the comment text to be processed belongs to the first type of starting part, the label "I-template 1" indicates that the current position in the comment text to be processed belongs to the first type of middle part, the label "B-template 2" indicates that the current position in the comment text to be processed belongs to the second type of starting part, and the label "I-template 2" indicates that the current position in the comment text to be processed belongs to the second type of middle part.
[0136] Optionally, in order to enrich the model, some samples that are not opinion phrases can be selected for training. These samples are labeled with a label "0", which indicates that the current position in the comment text to be processed is non-opinion information. In this case, if there are X templates, they can be labeled with (2*X+1) types of labels.
[0137] In the process of training to obtain the above sequence labeling model, the electronic device used for model training can pre-construct an initial sequence labeling model, and then input the above first sample into the initial sequence labeling model for training, thereby obtaining a sequence labeling model.
[0138] During the training process, the initial sequence labeling model can learn each first sample. After learning a large number of first samples, the initial sequence labeling model gradually establishes the correspondence between the opinion phrase and the starting position label and the middle position label, that is, learns how to use the starting position label and the middle position label to extract the opinion phrase from the text. In addition, the initial sequence labeling model can gradually establish the correspondence between the first position relationship and the second position relationship between the entity word and the opinion word in the opinion phrase based on the first position relationship between the starting position label and the middle position label, that is, learns how to use the position relationship between the starting position label and the middle position label to determine the position relationship between the entity word and the opinion word in the extracted opinion phrase. Among them, since the starting position label and the middle position label both include template identifiers, the entity words in the opinion phrase, the position relationship between the entity word and the opinion word in the opinion phrase, and the correspondence between the templates can be gradually established at the same time. Finally, the sequence labeling model is trained to obtain the sequence labeling model.
[0139] In the training process, when the initial sequence labeling model meets the first preset condition, the training can be stopped to obtain a trained sequence labeling model.
[0140] Optionally, the first preset condition may be that the number of iterations of each first sample reaches a preset number.
[0141] Optionally, the first preset condition may be that the error between the true value and the predicted value of the starting position label and the intermediate position label of each first sample is less than a preset error.
[0142] Among them, the sequence labeling model can adopt a deep learning model such as BERT (Bidirectional Encoder Representation from Transformers, a pre-trained language representation model) or a model with BiLSTM-CRF (Bi-directional Long Short-Term Memory-Conditional Random Field, a bidirectional long short-term memory network-conditional random field), and the embodiment of the present invention is not specifically limited to this. In addition, in the above-mentioned model with BiLSTM-CRF, BiLSTM is the abbreviation of Bi-directional Long Short-Term Memory (bidirectional long short-term memory network), BiLSTM is composed of forward LSTM (Long Short-Term Memory, long short-term memory network) and backward LSTM, and CRF is the abbreviation of Conditional Random Field (conditional random field).
[0143] In this way, after obtaining the comment text to be processed, the comment text to be processed can be directly input into the above sequence labeling model. The above sequence labeling model extracts the target opinion phrase from the comment text to be processed according to the established correspondence between the opinion phrase and the start position label and the middle position label, and the correspondence between the entity word in the opinion phrase and the template, and determines the target template corresponding to the target opinion phrase. Then, the above sequence labeling model can output the target opinion phrase and the target template corresponding to the target opinion phrase.
[0144] For example, Figure 2 As shown in the figure, Tok is the abbreviation of token (word), Tok1, Tok2...TokN are the input texts 1-N of the sequence labeling model, that is, the comment texts 1-N to be processed, E1, E2...EN are the vector representations (input embeddings) of the input texts 1-N respectively, T1, T2...TN are the encoded vectors (encodingvector) corresponding to the input texts 1-N respectively, and the entity words corresponding to each input text are obtained by connecting the classification layer to the encoded vector of each input text, thereby obtaining the target templates corresponding to each group of target opinion phrases and the target entity words in each group of target opinion phrases.
[0145] Optionally, in a specific implementation, based on the above step 41, in the above step S103, based on the correspondence between the description text and the opinion word preset for each entity word, selecting the target opinion word corresponding to the target description text from at least one candidate opinion word associated with the target placeholder in the target template may include the following step 51:
[0146] Step 51: Input the concatenated text obtained by concatenating the target opinion phrase and the target template into a preset opinion confirmation model, and obtain the target opinion word output by the opinion confirmation model.
[0147] Among them, the opinion confirmation model is obtained based on a preset plurality of second sample training; the second sample includes: a second sample opinion phrase and a second sample template corresponding to the second sample opinion phrase, the second sample template includes: a second sample entity word corresponding to the same entity as the entity word in the second sample opinion phrase, and a sample opinion word selected from at least one candidate opinion word associated with the second sample placeholder in the second sample template, corresponding to the description text in the second sample opinion phrase and used to replace the second sample placeholder in the second sample template.
[0148] In this specific implementation method, after obtaining the above-mentioned target opinion phrase and the above-mentioned target template, the above-mentioned target opinion phrase and the above-mentioned target template can be spliced, and then, the spliced text obtained by splicing is input into the preset opinion confirmation model, and the target opinion word output by the above-mentioned opinion confirmation model is obtained as the output result of the above-mentioned opinion confirmation model.
[0149] The opinion confirmation model is obtained through second sample training based on a plurality of preset second sample opinion phrases and second sample templates corresponding to the second sample opinion phrases, wherein the second sample templates include: second sample entity words corresponding to the same entity as the entity words in the above-mentioned second sample opinion phrases, and sample opinion words selected from at least one candidate opinion word associated with the second sample placeholder in the above-mentioned second sample template, corresponding to the description text in the above-mentioned second sample opinion phrase and used to replace the second sample placeholder in the above-mentioned second sample template.
[0150] In the process of training to obtain the above-mentioned opinion confirmation model, the electronic device used for model training can pre-construct an initial model, and then input the above-mentioned second sample into the initial model for training, thereby obtaining the opinion confirmation model.
[0151] During the training process, the initial model can learn each second sample. After learning a large number of second samples, the initial model gradually establishes the preset correspondence between the descriptive text and the opinion word for each entity word, and finally obtains the opinion confirmation model.
[0152] In the training process, when the initial sequence labeling model meets the second preset condition, the training can be stopped to obtain the opinion confirmation model.
[0153] Optionally, the second preset condition may be that the number of iterations of each second sample reaches a preset number.
[0154] Optionally, the second preset condition may be that the error between the true value and the predicted value of each second sample is less than a preset error.
[0155] In this way, after obtaining the concatenated text obtained by concatenating the target opinion phrase and the target template, the concatenated text can be directly input into the opinion confirmation model, thereby utilizing the correspondence between entity words and opinion words established in the opinion confirmation model, as well as the correspondence between description texts and opinion words preset for each entity word, to select the target opinion words from the candidate opinion words to replace the placeholder words in the target template as the output of the opinion confirmation model.
[0156] Among them, the opinion confirmation model can adopt a deep learning model with an attention mechanism (Attention Mehnism) such as BERT, which is not specifically limited in this embodiment of the present invention.
[0157] Optionally, in a specific implementation, the opinion confirmation model includes: a classification layer; the above step 51 may include the following steps 511-512:
[0158] Step 511: inputting the concatenated text obtained by concatenating the target opinion phrase and the target template into a preset opinion confirmation model to obtain a latent vector of the target placeholder determined based on the target opinion phrase and the specified content in the target template;
[0159] The specified content is: the content in the target template except the target placeholder;
[0160] Step 512: Input the latent vector to the classification layer so that the classification layer classifies the latent vector based on at least one candidate opinion word associated with the target placeholder word in the target template, and outputs the candidate opinion word representing the category to which the latent vector belongs as the target opinion word.
[0161] In this specific implementation, when training the above-mentioned opinion confirmation model, a correspondence between the description text in the target opinion phrase and the target opinion word used to replace the placeholder word in the above-mentioned target template can be established. Thus, when the above-mentioned opinion confirmation model receives the concatenated text obtained by concatenating the target opinion phrase and the target template, it can determine the latent vector of the target placeholder word in the above-mentioned target template based on the target opinion phrase and the specified content in the above-mentioned target template except the target placeholder word.
[0162] Furthermore, the determined latent vector can be input into the classification layer in the above-mentioned opinion confirmation model, and the above-mentioned classification layer can be used to classify the above-mentioned latent vector based on at least one candidate opinion word associated with the target placeholder in the above-mentioned target template. From the at least one candidate opinion word associated with the above-mentioned target placeholder, a candidate opinion word of the category to which the latent vector belongs is selected, and the selected candidate opinion word is output as the target opinion word.
[0163] Optionally, the target opinion phrase and the target template may be concatenated by identification to obtain a concatenated text.
[0164] For example, Figure 3 As shown, Figure 3 The [SEP] mark in the embodiment of the present invention is a mark for splicing the opinion phrase and the opinion template. Figure 3 [MASK] in the text is the identifier of the placeholder in the viewpoint template, which can be referred to as a placeholder. Figure 3 The [MASK] mark in is the target placeholder in the target template in the embodiment of the present invention, and Tokα, Tokβ, ..., Tokγ are the concatenated texts input into the opinion confirmation model, thus, Figure 3 The [CLS] marker in the figure is located before the starting position of the opinion phrase in the above concatenated text, that is, it is used to mark the complete input of the above opinion confirmation model. Eα, Eβ...Eγ are respectively the vector representations of the input texts Tokα, Tokβ...Tokγ, and Tα, Tβ...Tγ are respectively the encoded vectors corresponding to the input texts Tokα, Tokβ...Tokγ. By connecting the latent vector corresponding to the [MASK] placeholder (that is, the encoded vector) to the classification layer, the target opinion word corresponding to each input text is obtained.
[0165] The above-mentioned concatenated text is used as the input of the opinion confirmation model, and then the opinion confirmation model is used to learn the input concatenated text to obtain the latent vector of the target placeholder in the target template determined based on the above-mentioned target opinion phrase and the specified content in the target template. Then, the classification layer in the above-mentioned opinion confirmation model is used to classify the above-mentioned latent vector, and then the candidate opinion word of the category to which the above-mentioned latent vector belongs is obtained and output as the target opinion word.
[0166] For ease of understanding, Figure 4 FIG. 1 is a flow chart of a specific example of an embodiment of the present invention.
[0167] S401: performing data cleaning on the initial comment text, and performing specified processing on the initial comment text after data cleaning to obtain a comment text to be processed;
[0168] S402: Inputting the comment text to be processed into a preset sequence labeling model, and obtaining the target opinion phrase output by the sequence labeling model and the target template corresponding to the target opinion phrase;
[0169] S403: inputting the concatenated text obtained by concatenating the target opinion phrase and the target template into a preset opinion confirmation model to obtain a latent vector of the target placeholder determined based on the target opinion phrase and the specified content in the target template;
[0170] S404: Input the latent vector to the classification layer, so that the classification layer classifies the latent vector based on at least one candidate opinion word associated with the target placeholder word in the target template, and outputs the candidate opinion word representing the category to which the latent vector belongs as the target opinion word.
[0171] in, Figure 4 Step S401 in the embodiment of the present invention is step 21, Figure 4 Step S402 in the embodiment of the present invention is step 41. Figure 4 Step S403 in the embodiment of the present invention is step 511. Figure 4 Step S404 in is step 512 in the embodiment of the present invention and will not be described in detail here.
[0172] For example, Figure 5 As shown, Figure 5 The user comments in are the comment texts to be processed in the embodiment of the present invention. Figure 5 The vertical field cloze template library in the embodiment of the present invention is the viewpoint word template. Figure 5 The sequence labeling classification model in is the sequence labeling model in the embodiment of the present invention. Figure 5 The opinion phrase and its corresponding template in the embodiment of the present invention are the spliced text obtained by splicing the target opinion phrase and the target template. Figure 5 The [MASK] classification model in the embodiment of the present invention is the opinion confirmation model. Figure 5 The normalized viewpoint in is the target viewpoint in the embodiment of the present invention.
[0173] The collected user comments and vertical field cloze template library are input into the sequence labeling classification model to obtain opinion phrases and their corresponding templates. Then, the concatenated text obtained by concatenating the above opinion phrases and their corresponding templates is input into the [MASK] classification model, and the output result of the above [MASK] classification model is used as the normalized opinion.
[0174] Corresponding to the text processing method provided by the above-mentioned embodiment of the present invention, the embodiment of the present invention also provides a text processing device.
[0175] Figure 6 A structural diagram of a text processing device provided by an embodiment of the present invention is shown in FIG. Figure 6 As shown, the device may include the following modules:
[0176] The text acquisition module 610 is used to acquire the comment text to be processed;
[0177] The target template determination module 620 is used to extract the target opinion phrase from the comment text to be processed, and based on the positional relationship between the target entity word in the target opinion phrase and the target description text, select the target template corresponding to the target entity word in the target opinion phrase from the preset opinion word template; wherein the target description text is used to describe the entity corresponding to the target entity word, each opinion word template includes: an entity word and a placeholder, and at least one candidate opinion word associated with the placeholder, and the candidate opinion word included in each opinion word template is used to describe the entity corresponding to the entity word included in the opinion word template;
[0178] A target opinion word determination module 630 is used to select a target opinion word corresponding to the target description text from at least one candidate opinion word associated with the target placeholder in the target template based on the correspondence relationship between the description text and the opinion word preset for each entity word;
[0179] The result acquisition module 640 is used to replace the target placeholder word with the target opinion word, and determine the replaced target template as the processing result of the comment text to be processed.
[0180] As can be seen from the above, when the scheme provided by the embodiment of the present invention is applied, when processing the text, first, the comment text to be processed is obtained, and then the target opinion phrase is extracted from the comment text to be processed, and based on the positional relationship between the target entity word and the target description text in the target opinion phrase, the target template corresponding to the target entity word is selected from the preset opinion word template. Since each opinion word template includes: an entity word and a placeholder, and the placeholder is associated with at least one candidate opinion word for describing the entity corresponding to the entity word included in the opinion word template, the selected target template includes: a target entity word and a target placeholder, and the target placeholder is associated with at least one candidate opinion word. Then, based on the preset correspondence relationship between the description text and the opinion word for each entity word, the target opinion word corresponding to the target description text can be selected from at least one candidate opinion word associated with the target placeholder in the selected target template, so that the target opinion word can be used to replace the target placeholder, and the replaced target template is determined as the processing result of the comment text to be processed.
[0181] Based on this, by applying the solution provided by the embodiment of the present invention, with the help of a preset opinion word template and the preset correspondence between the description text and the opinion word for each entity word, the description text in the target opinion phrase for the same entity extracted from the above-mentioned text to be processed can be unified into the same opinion word, thereby realizing normalization of the description text, making it easier for users to intuitively obtain valid comment information about the entity based on the opinion words and perform subsequent operations on the above-mentioned valid comment information, thereby improving the intuitiveness of the text and improving the user's reading experience of the extracted opinion phrases.
[0182] Optionally, in a specific implementation, the target template determination module 620 is specifically used to:
[0183] Inputting the review text to be processed into a preset sequence labeling model, and obtaining a target opinion phrase output by the sequence labeling model and a target template corresponding to the target opinion phrase;
[0184] Among them, the sequence labeling model is obtained based on the training of multiple preset first samples; the first sample includes: a first sample opinion phrase marked with a starting position label and an intermediate position label, the starting position label includes: a starting identifier for representing the beginning of the first sample opinion phrase and a sample template identifier of the first sample template corresponding to the first sample opinion phrase; the intermediate position label includes: an intermediate label for representing the end of the first sample opinion phrase and the sample template identifier.
[0185] Optionally, in a specific implementation, the target opinion word determination module 630 includes:
[0186] A target opinion word acquisition submodule is used to input the concatenated text obtained by concatenating the target opinion phrase and the target template into a preset opinion confirmation model, and obtain the target opinion word output by the opinion confirmation model;
[0187] Wherein, the opinion confirmation model is obtained by training based on a plurality of preset second samples; the second sample includes: a second sample opinion phrase and a second sample template corresponding to the second sample opinion phrase, the second sample template includes: a second sample entity word corresponding to the same entity as the entity word in the second sample opinion phrase, and a sample opinion word selected from at least one candidate opinion word associated with the second sample placeholder in the second sample template, corresponding to the description text in the second sample opinion phrase and used to replace the second sample placeholder.
[0188] Optionally, in a specific implementation, the opinion confirmation model includes: a classification layer; the target opinion word acquisition submodule is specifically used for:
[0189] Inputting the concatenated text obtained by concatenating the target opinion phrase and the target template into a preset opinion confirmation model, obtaining a latent vector of the placeholder in the target template determined based on the target opinion phrase and the specified content in the target template; wherein the specified content is: the content in the target template other than the target placeholder;
[0190] The latent vector is input to the classification layer so that the classification layer classifies the latent vector based on at least one candidate opinion word associated with the target placeholder in the target template, and outputs the candidate opinion word representing the category to which the latent vector belongs as the target opinion word.
[0191] Optionally, in a specific implementation, the device further includes:
[0192] The template acquisition module is used to obtain initial opinion phrases whose appearance frequency in the comment text to be analyzed is higher than a specified frequency; replace the opinion words in the initial opinion phrases with preset placeholders, and determine the replaced opinion words as candidate opinion words associated with the placeholders to obtain an initial template; merge the initial templates for the same entity and including the same positional relationship between the entity words and the placeholders to obtain individual opinion word templates.
[0193] Optionally, in a specific implementation, the text acquisition module 610 includes:
[0194] The initial text acquisition submodule is used to obtain the initial comment text;
[0195] The initialization submodule is used to perform data initialization on the initial comment text to obtain the comment text to be processed.
[0196] Optionally, in a specific implementation, the initialization submodule is specifically used to:
[0197] Performing data cleaning on the initial comment text, and performing specified processing on the initial comment text after data cleaning to obtain the comment text to be processed;
[0198] The designated processing includes: word segmentation processing and / or character segmentation processing.
[0199] Corresponding to the text processing method provided by the above embodiment of the present invention, the embodiment of the present invention also provides an electronic device, such as Figure 7 As shown, it includes a processor 701, a communication interface 702, a memory 703 and a communication bus 704, wherein the processor 701, the communication interface 702, and the memory 703 communicate with each other through the communication bus 704.
[0200] Memory 703, used for storing computer programs;
[0201] The processor 701 is used to implement the steps of any text processing method provided by the above-mentioned embodiments of the present invention when executing the program stored in the memory 703.
[0202] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0203] The communication interface is used for communication between the above electronic device and other devices.
[0204] The memory may include a random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0205] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0206] In another embodiment of the present invention, a computer-readable storage medium is provided, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned text processing methods are implemented.
[0207] In another embodiment of the present invention, a computer program product including instructions is provided, which, when executed on a computer, enables the computer to execute the steps of any text processing method in the above embodiments.
[0208] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk Solid State Disk (SSD)), etc.
[0209] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0210] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, mobile robot embodiment, computer-readable storage medium embodiment, and computer program product embodiment, since they are basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0211] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A text processing method, It is characterized in that The method comprises: Get the comment text to be processed; A target opinion phrase is extracted from the comment text to be processed, and based on the positional relationship between the target entity word in the target opinion phrase and the target description text, a target template corresponding to the target entity word in the target opinion phrase is selected from the preset opinion word template; wherein the target description text is used to describe the entity corresponding to the target entity word, each opinion word template includes: an entity word and a placeholder word, and the placeholder word is associated with at least one candidate opinion word, and the candidate opinion word included in each opinion word template is used to describe the entity corresponding to the entity word included in the opinion word template; the entity word in the selected target template corresponds to the same entity as the target entity word in the target opinion phrase, and the positional relationship between the entity word and the placeholder included in the selected target template is the same as the positional relationship between the target entity word and the target description text included in the target opinion phrase; Based on the correspondence between the description text and the opinion word preset for each entity word, a target opinion word corresponding to the target description text is selected from at least one candidate opinion word associated with the target placeholder word in the target template; The target placeholder is replaced by the target opinion word, and the replaced target template is determined as a processing result of the comment text to be processed.
2. The method according to claim 1, It is characterized in that The step of extracting a target opinion phrase from the review text to be processed, and selecting a target template corresponding to the target entity word in the target opinion phrase from a preset opinion word template based on the positional relationship between the target entity word in the target opinion phrase and the target description text, includes: Inputting the review text to be processed into a preset sequence labeling model, and obtaining a target opinion phrase output by the sequence labeling model and a target template corresponding to the target opinion phrase; Among them, the sequence labeling model is obtained based on the training of multiple preset first samples; the first sample includes: a first sample opinion phrase marked with a starting position label and an intermediate position label, the starting position label includes: a starting identifier for representing the beginning of the first sample opinion phrase and a sample template identifier of the first sample template corresponding to the first sample opinion phrase; the intermediate position label includes: an intermediate label for representing the end of the first sample opinion phrase and the sample template identifier.
3. The method according to claim 2, It is characterized in that The selecting, based on the correspondence between the description text and the opinion word preset for each entity word, a target opinion word corresponding to the target description text from at least one candidate opinion word associated with the target placeholder word in the target template comprises: Inputting the target opinion phrase and the target template into a concatenated text into a preset opinion confirmation model, and obtaining the target opinion word output by the opinion confirmation model; Wherein, the opinion confirmation model is obtained by training based on a plurality of preset second samples; the second sample includes: a second sample opinion phrase and a second sample template corresponding to the second sample opinion phrase, the second sample template includes: a second sample entity word corresponding to the same entity as the entity word in the second sample opinion phrase, and a sample opinion word selected from at least one candidate opinion word associated with the second sample placeholder in the second sample template, corresponding to the description text in the second sample opinion phrase and used to replace the second sample placeholder.
4. The method according to claim 3, It is characterized in that The opinion confirmation model includes: a classification layer; the step of inputting the spliced text obtained by splicing the target opinion phrase and the target template into a preset opinion confirmation model, and obtaining the target opinion word output by the opinion confirmation model, including: Inputting the concatenated text obtained by concatenating the target opinion phrase and the target template into a preset opinion confirmation model to obtain a latent vector of the target placeholder determined based on the target opinion phrase and the specified content in the target template; wherein the specified content is: the content in the target template other than the target placeholder; The latent vector is input to the classification layer so that the classification layer classifies the latent vector based on at least one candidate opinion word associated with the target placeholder in the target template, and outputs the candidate opinion word representing the category to which the latent vector belongs as the target opinion word.
5. The method according to any one of claims 1 to 4, It is characterized in that The method of constructing the opinion word template includes: Obtaining initial opinion phrases that appear in the comment text to be analyzed at a frequency higher than a specified frequency; Replacing the opinion word in the initial opinion phrase with a preset placeholder, and determining the replaced opinion word as a candidate opinion word associated with the placeholder, to obtain an initial template; The initial templates for the same entity and including entity words and placeholder words with the same positional relationship are merged to obtain various opinion word templates.
6. The method according to any one of claims 1 to 4, It is characterized in that The step of obtaining the comment text to be processed includes: An initial comment text is obtained, and data of the initial comment text is initialized to obtain a comment text to be processed.
7. The method according to claim 6, It is characterized in that The step of initializing the data of the initial comment text to obtain the comment text to be processed includes: The initial comment text is cleaned and the initial comment text after data cleaning is subjected to designated processing to obtain the comment text to be processed; wherein the designated processing includes: word segmentation processing and / or character segmentation processing.
8. An electronic device, It is characterized in that It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing the method steps described in any one of claims 1 to 7 when executing a program stored in a memory.
9. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Semantic-based methods and search systems for finding, integrating, and providing comment information.
CN102279894A
Comment information processing method and device, computer equipment and medium
CN111191428A