A method and device for determining an outline
By extracting the semantic features of the text to be generated and matching it with the semantic features of the preset outline, the outline is automatically selected, and the problem of inefficient outline determination in the prior art is solved, and more efficient outline selection is achieved.
Patent Information
- Application Number
- CN202110880841.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-02
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-08-02
AI Technical Summary
In the prior art, texts need to be manually selected from multiple outlines before automatically generating text, resulting in low efficiency in determining outlines.
By obtaining the description information of the text to be generated, its semantic features are extracted, and based on the similarity between the semantic features of the preset outline and the first semantic features, the outline of the text to be generated is automatically selected from the preset outline.
It improves the efficiency of outline determination, and can accurately select suitable outlines from preset outlines, reducing the time and labor force for manual selection.
Smart Images

Figure CN113688633B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to an outline determination method and apparatus. Background Art
[0002] Automatic text generation is a research branch of natural language processing, which enables an electronic device to generate text. The text generated by the electronic device can assist users in efficiently creating high-quality text. Before generating the text, it is usually necessary to determine the outline of the text, and the electronic device generates the text based on each outline.
[0003] In the prior art, usually, a staff member manually selects an outline from a relatively large number of outlines. However, this method is time-consuming and laborious, resulting in low efficiency in outline determination. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide an outline determination method and apparatus to improve the efficiency of outline determination. The specific technical solutions are as follows:
[0005] In a first aspect, the embodiments of the present invention provide an outline determination method, and the method includes:
[0006] Obtain the description information of the text to be generated;
[0007] Obtain the semantic feature of the description information as the first semantic feature;
[0008] Based on the semantic feature of the preset outline and the first semantic feature, select the outline of the text to be generated from the preset outline.
[0009] In an embodiment of the present invention, the above-mentioned determining the outline of the text to be generated from the preset outline based on the semantic feature of the preset outline and the first semantic feature includes:
[0010] Based on the similarity between the semantic feature of the preset outline and the first semantic feature, select the outline of the text to be generated from the preset outline.
[0011] In an embodiment of the present invention, the above-mentioned selecting the outline of the text to be generated from the preset outline based on the similarity between the semantic feature of the preset outline and the first semantic feature includes:
[0012] Based on the similarity between the semantic feature of the clustering center of each outline group and the first semantic feature, select the outline group to which the outline of the text to be generated belongs from each outline group as an alternative outline group, where each outline group is an outline group obtained by clustering according to the similarity between the semantic features of the outlines;
[0013] Select the outline of the text to be generated from each outline in the alternative outline group according to the similarity between the semantic features of each outline in the alternative outline group and the first semantic feature.
[0014] In one embodiment of the present invention, the step of selecting the outline of the text to be generated from each outline in the alternative outline group according to the similarity between the semantic features of each outline in the alternative outline group and the first semantic feature includes:
[0015] Select from the alternative outline group the alternative outline group whose semantic feature of the cluster center has the highest similarity with the first semantic feature;
[0016] Select the outline of the text to be generated from each outline in the selected alternative outline group according to the similarity between the semantic features of each outline in the selected alternative outline group and the first semantic feature.
[0017] In one embodiment of the present invention, the step of selecting the outline of the text to be generated from each outline in the alternative outline group according to the similarity between the semantic features of each outline in the alternative outline group and the first semantic feature includes:
[0018] Calculate the similarity between the semantic features of each outline in the alternative outline group and the first semantic feature;
[0019] Select the first preset number of outlines from each outline in descending order of the calculated similarity corresponding to each outline as the outline of the text to be generated.
[0020] In one embodiment of the present invention, the step of selecting the outline group to which the outline of the text to be generated belongs from each outline group based on the similarity between the semantic features of the cluster centers of each outline group and the first semantic feature includes:
[0021] Calculate the similarity between the semantic features of the cluster centers of each outline group and the first semantic feature;
[0022] Select the first second preset number of outline groups from each outline group in descending order of the calculated similarity corresponding to each outline group;
[0023] Determine the number of outlines containing the description information in each of the selected outline groups;
[0024] Determine the outline group to which the outline of the text to be generated belongs from the selected outline groups according to the determined number of outlines.
[0025] In one embodiment of the present invention, the method further includes:
[0026] Select the paragraphs of the text to be generated from the preset paragraphs corresponding to the outline of the text to be generated.
[0027] In one embodiment of the present invention, the above-mentioned selection of the paragraphs of the text to be generated from the preset paragraphs corresponding to the outline of the text to be generated includes:
[0028] Based on the similarity between the semantic features of the preset paragraphs corresponding to the outline of the text to be generated and the first semantic features, select the paragraphs of the text to be generated from the preset paragraphs corresponding to the outline of the text to be generated.
[0029] In one embodiment of the present invention, the above-mentioned preset paragraphs are pre-determined paragraphs, and the following steps are included:
[0030] Obtain the pre-selected text corresponding to the preset outline;
[0031] Extract paragraphs from each paragraph corresponding to the preset outline in the text as the preset paragraphs corresponding to the outline.
[0032] In one embodiment of the present invention, the above-mentioned extraction of paragraphs from each paragraph corresponding to the preset outline in the text as the preset paragraphs corresponding to the outline includes:
[0033] Determine the feature information of each paragraph corresponding to the preset outline in the text;
[0034] Based on the feature information of each paragraph, select alternative paragraphs from each paragraph, and determine the alternative paragraphs as the preset paragraphs corresponding to the preset outline in the text.
[0035] In one embodiment of the present invention, the above-mentioned determination of the alternative paragraphs as the preset paragraphs corresponding to the preset outline in the text includes:
[0036] For each alternative paragraph, determine the semantic features of the alternative paragraph and the semantic features of each word in the alternative paragraph, and input the determined semantic features and semantic features into a pre-trained paragraph quality evaluation model to obtain the quality score value of the alternative paragraph. Select the paragraphs with quality score values greater than the preset quality score threshold as the preset paragraphs corresponding to the preset outline in the text;
[0037] Among them, the paragraph quality evaluation model is: a neural network model preset is trained with the semantic features of the sample paragraphs and the semantic features of each word in the sample paragraphs as the model input and the labeled quality score value of the sample paragraphs as the training benchmark to obtain the quality score value of the paragraph.
[0038] In one embodiment of the present invention, the above method further includes:
[0039] Based on the outline of the text to be generated, sort the selected paragraphs of the text to be generated, and generate a text containing the outline of the text to be generated and the sorted paragraphs.
[0040] In one embodiment of the present invention, the above-described description information includes at least one of the following information: user portrait, keyword, entity word, key sentence, and text type.
[0041] In a second aspect, an embodiment of the present invention provides an outline determination device, and the device includes:
[0042] An information acquisition module, configured to acquire description information of the text to be generated;
[0043] A feature acquisition module, configured to acquire a semantic feature of the description information as a first semantic feature;
[0044] An outline selection module, configured to select an outline of the text to be generated from the preset outlines based on the semantic feature of the preset outline and the first semantic feature.
[0045] In one embodiment of the present invention, the above-described outline selection module is specifically configured to select an outline of the text to be generated from the preset outlines based on the similarity between the semantic feature of the preset outline and the first semantic feature.
[0046] In one embodiment of the present invention, the above-described outline selection module includes:
[0047] An outline group selection sub-module, configured to select, as an alternative outline group, an outline group to which the outline of the text to be generated belongs from each outline group based on the similarity between the semantic feature of the clustering center of each outline group and the first semantic feature, where each outline group is an outline group obtained by clustering according to the similarity between semantic features of outlines;
[0048] An outline selection sub-module, configured to select an outline of the text to be generated from each outline in the alternative outline group according to the similarity between the semantic feature of each outline in the alternative outline group and the first semantic feature.
[0049] In one embodiment of the present invention, the above-described outline selection sub-module is specifically configured to select, from the alternative outline groups, an alternative outline group with the highest similarity between the semantic feature of the clustering center and the first semantic feature; and select an outline of the text to be generated from each outline within the selected alternative outline group according to the similarity between the semantic feature of each outline in the selected alternative outline group and the first semantic feature.
[0050] In one embodiment of the present invention, the above-mentioned outline selection sub-module is specifically configured to calculate the similarity between the semantic features of each outline in the alternative outline group and the first semantic feature; select the first preset number of outlines from each outline in descending order of the calculated similarity corresponding to each outline as the outline of the text to be generated.
[0051] In one embodiment of the present invention, the above-mentioned outline group selection sub-module includes:
[0052] A similarity calculation unit, configured to calculate the similarity between the semantic features of the clustering centers of each outline group and the first semantic feature;
[0053] An outline group selection unit, configured to select the first second preset number of outline groups from each outline group in descending order of the calculated similarity corresponding to each outline group;
[0054] A quantity determination unit, configured to determine the number of outlines in the selected outline groups that contain the description information;
[0055] An outline group determination unit, configured to determine the outline group to which the outline of the text to be generated belongs from the selected outline groups according to the determined number of outlines.
[0056] In one embodiment of the present invention, the above-mentioned device further includes a paragraph selection module,
[0057] The paragraph selection module is specifically configured to select the paragraphs of the text to be generated from the preset paragraphs corresponding to the outline of the text to be generated.
[0058] In one embodiment of the present invention, the above-mentioned paragraph selection module is specifically configured to select the paragraphs of the text to be generated from the preset paragraphs corresponding to the outline of the text to be generated based on the similarity between the semantic features of the preset paragraphs corresponding to the outline of the text to be generated and the first semantic feature.
[0059] In one embodiment of the present invention, the above-mentioned device further includes a preset paragraph determination module, and the preset paragraph determination module includes:
[0060] A text acquisition sub-module, configured to acquire the pre-selected text corresponding to the preset outline;
[0061] A paragraph determination sub-module, configured to extract paragraphs from the paragraphs corresponding to the preset outline in the text as the preset paragraphs corresponding to the outline.
[0062] In one embodiment of the present invention, the above-mentioned paragraph determination sub-module includes:
[0063] An information determination unit, configured to determine the feature information of the paragraphs corresponding to the preset outline in the text;
[0064] A paragraph determination unit, configured to select alternative paragraphs from the respective paragraphs based on the feature information of the respective paragraphs, and determine the alternative paragraphs as the preset paragraphs corresponding to the preset outline in the text.
[0065] In one embodiment of the present invention, the above-mentioned paragraph determination unit is specifically configured to, for each alternative paragraph, determine the semantic feature of the alternative paragraph and the semantic feature of each word in the alternative paragraph, and input the determined semantic feature and semantic feature into a pre-trained paragraph quality evaluation model to obtain a quality score value of the alternative paragraph, and use the paragraph whose quality score value is greater than a preset quality score threshold as the preset paragraph corresponding to the preset outline in the text;
[0066] Wherein, the paragraph quality evaluation model is: obtained by training a preset neural network model with the semantic features of the sample paragraphs and the semantic features of each word in the sample paragraphs as the model input and the labeled quality score value of the sample paragraphs as the training benchmark, and is used to obtain the quality score value of the paragraph.
[0067] In one embodiment of the present invention, the above-mentioned device further includes: a text generation module,
[0068] The text generation module is specifically configured to sort the paragraphs of the text to be generated selected based on the outline of the text to be generated, and generate a text including the outline of the text to be generated and the sorted paragraphs.
[0069] In one embodiment of the present invention, the above-mentioned description information includes at least one of the following information: user portrait, keyword, entity word, key sentence, and text type.
[0070] In a third aspect, an embodiment of the present invention provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0071] The memory is used to store a computer program;
[0072] The processor is configured to implement the method steps described in the first aspect when executing the program stored on the memory.
[0073] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the method steps described in the first aspect are implemented.
[0074] As can be seen from the above, when using the solution provided by the embodiments of the present invention to determine an outline, since the outline of the text to be generated is determined from the preset outline based on the semantic features of the preset outline and the first semantic features, compared with the prior art, it is not necessary for staff to manually determine the outline, which improves the efficiency of outline determination.
[0075] In addition, since the above-mentioned first semantic features can reflect the semantics expressed by the description information of the text to be generated, and the semantic features of the preset outline can reflect the semantics expressed by each preset outline, based on these two types of information, the outline of the text to be generated can be determined from the preset outline more accurately.
[0076] Of course, it is not necessary for any product or method implementing the present invention to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0078] Figure 1 It is a schematic flowchart of the first outline determination method provided by the embodiments of the present invention;
[0079] Figure 2 It is a schematic flowchart of an outline selection method provided by the embodiments of the present invention;
[0080] Figure 3 It is a schematic flowchart of the second outline determination method provided by the embodiments of the present invention;
[0081] Figure 4 It is a schematic flowchart of the third outline determination method provided by the embodiments of the present invention;
[0082] Figure 5 It is a flowchart of a method for obtaining paragraphs provided by the embodiments of the present invention;
[0083] Figure 6 It is a flowchart of a method for obtaining text information provided by the embodiments of the present invention;
[0084] Figure 7 It is a flowchart of a process for calculating semantic similarity based on a model provided by the embodiments of the present invention;
[0085] Figure 8 It is a flowchart of a text generation method provided by the embodiments of the present invention;
[0086] Figure 9a A flowchart of a process for obtaining a preset paragraph provided by an embodiment of the present invention;
[0087] Figure 9b A flowchart of a process for evaluating the quality of alternative paragraphs provided by an embodiment of the present invention;
[0088] Figure 9c A flowchart of an Attention mechanism provided by an embodiment of the present invention;
[0089] Figure 10 A structural diagram of a first outline determination device provided by an embodiment of the present invention;
[0090] Figure 11 A structural diagram of an outline selection module provided by an embodiment of the present invention;
[0091] Figure 12 A structural diagram of a second outline determination device provided by an embodiment of the present invention;
[0092] Figure 13 A structural diagram of a third outline determination device provided by an embodiment of the present invention;
[0093] Figure 14 A structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0094] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0095] See Figure 1 , Figure 1 A flowchart of a first outline determination method provided by an embodiment of the present invention, and the above method includes S101 - S103.
[0096] The execution subject of the embodiment of the present invention may be an electronic device, such as a server, a laptop computer, etc.
[0097] S101: Obtain the description information of the text to be generated.
[0098] The text to be generated may be texts such as news, articles, essays, etc.
[0099] The description information of the text to be generated can be understood as: information used to describe the text to be generated. For example, the description information of the text to be generated may include keywords, titles, text types, etc. of the text to be generated.
[0100] The description information is for describing the information of the text to be generated, and to a certain extent, it can reflect the basic characteristics of the text to be generated, such as characteristics like the type, style, keywords, etc. of the text to be generated. The description information can also be called feature information.
[0101] Specifically, the information input by the user for describing the text to be generated can be used as the description information of the text to be generated.
[0102] In one implementation, the user can directly input the description information of the text to be generated.
[0103] For example: The user inputs keywords, types, titles, etc. of the text to be generated, and the electronic device uses the information input by the user as the description information of the text to be generated.
[0104] In another implementation, the user can also input a description text segment of the text to be generated, and the electronic device extracts the description information from the above description text segment to obtain the description information of the text to be generated.
[0105] Specifically, after obtaining the description text segment, the electronic device can clean and filter the description text segment. For example: It can perform processing such as sensitive word filtering and stop word filtering on the above description text segment, and then extract the description information from the cleaned and filtered description text segment.
[0106] In yet another implementation, the electronic device can also obtain a user profile and use both the description information input by the user and the user profile as the description information of the text to be generated.
[0107] Specifically, when obtaining the above user profile, it can be based on the user identifier and the preset correspondence between the user identifier and the user profile to obtain the user profile. The above user identifier can be the user's ID number, login name, etc.
[0108] The above user profile can be understood as information used to describe the user. The above user profile can include description information in multiple different dimensions. For example: The above user profile can include the user's attribute information, interest information, behavior information, scenario information, etc.
[0109] S102: Obtain the semantic features of the description information as the first semantic features.
[0110] Semantic features are used to reflect the semantics expressed by an object. In the embodiments of the present invention, semantic features of different objects appear. For each object, the semantic features of the object are used to reflect the semantics expressed by the object.
[0111] The semantic features of the description information are used to reflect the semantics expressed by the description information. When obtaining the semantic features of the description information, semantic features represented in a vectorized form can be obtained. For example: The description information can be vectorized and encoded, and the encoded result can be used as the semantic features of the description information. It is also possible to analyze the semantic information of the description information and extract semantic features based on the analysis results.
[0112] Step S103: Determine the outline of the text to be generated from the preset outline based on the semantic features of the preset outline and the first semantic features.
[0113] An outline is used to reflect the structural information of the text. Specifically, the outline can be the various subheadings in the text, and the outline can also be the central sentences of each paragraph in the text.
[0114] For example: Taking a patent text as an example, the outline of the patent text can include: abstract of the specification, description of the drawings, claims, specification, drawings of the specification; taking an academic paper as an example, the outline of the academic paper can include: abstract, keywords, specific content, references.
[0115] Since the various subheadings of the text are usually summary headings, and the central sentences of each paragraph in the text are usually used to summarize the central idea of the paragraph, and the text is usually formed based on the central idea of each paragraph or the content of each subheading in a certain logical order, the outline can more accurately reflect the structural information of the text.
[0116] The preset outline can be an outline extracted from various types of texts obtained in advance. Specifically, the electronic device can store historical texts, extract the outlines in the stored texts, and use the extracted outlines as the preset outlines; the electronic device can also, based on an automatic crawler system, regularly and automatically crawl the texts of a specified website and save them in a database, thereby obtaining the incremental texts of the specified website, and conduct real-time monitoring of the text information on the Internet, discover the sources of text data, detect the dynamic information of text data, so as to obtain a relatively large number of texts, extract the outlines in the obtained texts, and save the extracted outlines of the texts as the preset outlines.
[0117] The semantic features of the preset outline are used to reflect the semantics expressed by the preset outline. Specifically, the semantic features of the preset outline can be that the electronic device pre-vectorizes and encodes each preset outline, and uses the encoded vectors as the semantic features of each preset outline.
[0118] In an embodiment of the present invention, the outline of the text to be generated can be selected from the preset outline based on the similarity between the semantic features of the preset outline and the first semantic features.
[0119] The similarity between the semantic features of the above-mentioned preset outline and the first semantic feature can be determined by calculating the distance between the semantic features of the preset outline and the first semantic feature, and based on the calculated distance. The above distance can be Euclidean distance, cosine distance, etc. For example: For example: A distance similarity conversion algorithm can be used to convert the calculated distance into a similarity.
[0120] When selecting the outline of the text to be generated from the preset outline based on the similarity between the semantic features of the preset outline and the first semantic feature, it can be to select a preset number of preset outlines with the highest similarity as the outline of the text to be generated, or it can be to select the preset outlines with a similarity greater than the preset similarity as the outline of the text to be generated.
[0121] The above preset number can be set by the staff according to experience. For example: The above preset number can be 1, 2, 3, 4, etc. Taking the first preset number as 3 as an example, select 3 preset outlines with the highest similarity as the outline of the text to be generated.
[0122] In this way, since the outline of the text to be generated is determined based on the similarity between the semantic features of the preset outline and the first semantic feature, and the above first semantic feature can reflect the semantics expressed by the description information of the text to be generated, and the semantic features of the preset outline can reflect the semantics expressed by each preset outline, therefore, based on the similarity between the above semantic features, the outline of the text to be generated can be determined more accurately.
[0123] As can be seen from the above, when using the solution provided in this embodiment to determine the outline, since the outline of the text to be generated is determined from the preset outline based on the semantic features of the preset outline and the first semantic feature, compared with the prior art, it does not require the staff to manually determine the outline, improving the efficiency of outline determination.
[0124] In addition, since the above first semantic feature can reflect the semantics expressed by the description information of the text to be generated, and the semantic features of the preset outline can reflect the semantics expressed by each preset outline, based on the above two types of information, the outline of the text to be generated can be determined more accurately from the preset outline.
[0125] In an embodiment of the present invention, the above preset outline can be; the outlines included in multiple outline groups obtained by clustering according to the similarity between the semantic features of the outlines.
[0126] Specifically, based on the similarity between the semantic features of each preset outline, the semantic features of each preset outline can be clustered. For example: The K-means (k-means clustering) algorithm can be used to cluster the semantic features of each preset outline to obtain the clustered outline groups.
[0127] After determining the outline groups after each clustering, the clustering centers in each outline group can be determined. The clustering center refers to an outline with relatively high similarity to each of the preset outlines included in the outline group.
[0128] Specifically, a distributed file storage system can be used to store the semantic feature vector data of each outline in each clustered outline group and the semantic feature vector data of the clustering center, and the semantic feature vector data of the clustering center of each clustered outline group is saved in the memory of the electronic device.
[0129] Based on the above embodiments, refer to Figure 2 , Figure 2 which is a schematic flowchart of a method for selecting an outline provided by an embodiment of the present invention. The above step S103 of selecting an outline for the text to be generated from the preset outlines according to the similarity between the semantic feature of the preset outline and the first semantic feature can be implemented according to the following steps S201-S202.
[0130] S201: Based on the similarity between the semantic feature of the clustering center of each outline group and the first semantic feature, select the outline group to which the outline for the text to be generated belongs from each outline group as an alternative outline group.
[0131] The above-mentioned each outline group is: an outline group obtained by clustering according to the similarity between the semantic features of the outlines.
[0132] Since the above-mentioned outline group is obtained by clustering according to the similarity between the semantic features of the outlines, each outline group has a clustering center. The clustering center refers to an outline with relatively high similarity to each of the preset outlines included in the outline group, and the semantic feature of the clustering center can reflect the overall semantic feature of each outline group.
[0133] The above-mentioned first semantic feature is: the semantic feature of the description information.
[0134] Specifically, since the semantic features of the clustering centers of the above-mentioned each outline group can be stored in the memory of the electronic device, when calculating the similarity between the semantic features of the clustering centers of each outline group and the first semantic feature, the semantic features of the clustering centers of each outline group can be obtained from the memory of the electronic device, thereby improving the obtaining efficiency of the semantic features of the clustering centers.
[0135] When determining the similarity between the semantic feature of the clustering center of each outline group and the first semantic feature, it can be to calculate the distance between the semantic feature of the preset clustering center of each outline group and the above-mentioned first semantic feature, and determine the above-mentioned similarity based on the calculated distance. The above-mentioned distance can be Euclidean distance, cosine distance, etc. For example: a distance similarity conversion algorithm can be used to convert the calculated distance into similarity.
[0136] When selecting the outline group to which the outline of the text to be generated belongs from each outline group as the alternative outline group based on the similarity between the semantic features of the clustering center of each outline group and the first semantic feature, it can be to select a preset number of outline groups with the highest similarity as the alternative outline group, or it can be to select the outline groups with similarity greater than the preset similarity as the alternative outline group.
[0137] S202: Select the outline of the text to be generated from each outline in the alternative outline group according to the similarity between the semantic features of each outline in the alternative outline group and the first semantic feature.
[0138] Specifically, when determining the similarity between the semantic features of each outline in the alternative outline group and the first semantic feature, the distance between the semantic features of each outline and the above first semantic feature can be calculated, and the above similarity can be determined based on the calculated distance. The above distance can be Euclidean distance, cosine distance, etc. For example: The distance similarity conversion algorithm can be used to convert the calculated distance into similarity.
[0139] Specifically, when selecting the outline of the text to be generated from each outline in the alternative outline group according to the similarity between the semantic features of each outline in the alternative outline group and the first semantic feature, it can be to select a preset number of outlines with the highest similarity, or it can be to select the outlines with similarity greater than the preset similarity.
[0140] In this way, since first the outline group to which the outline of the text to be generated belongs is selected from each clustered outline group as the alternative outline group; then the outline of the text to be generated is selected from each outline in the alternative outline group. And the clustering of each outline group is based on the similarity between the semantic features of each outline, that is, each relatively similar outline is divided into an outline group. Compared with selecting the outline of the text to be generated with each outline as a unit, the efficiency of selecting the outline of the text to be generated with the outline group as a unit is higher.
[0141] In an embodiment of the present invention, the above S202 of selecting the outline of the text to be generated from each outline in the alternative outline group according to the similarity between the semantic features of each outline in the alternative outline group and the first semantic feature can be implemented according to the following steps A1 - A2.
[0142] Step A1: Select the alternative outline group with the highest similarity between the semantic features of the clustering center and the first semantic feature from the alternative outline groups.
[0143] Specifically, based on the similarity between the semantic features of the clustering center of each alternative outline group and the first semantic feature, the alternative outline group with the highest similarity can be selected from each alternative outline group.
[0144] Step A2: Select the outline for the text to be generated from the outlines within the selected alternative outline group according to the similarity between the semantic features of each outline in the selected alternative outline group and the first semantic feature.
[0145] Specifically, when selecting the outline for the text to be generated, a preset number of outlines with the highest similarity can be selected as the outline for the text to be generated. Outlines with a similarity greater than the preset similarity can also be selected as the outline for the text to be generated.
[0146] In this way, since the semantic features of the clustering center of the selected alternative outline group have the highest similarity with the first semantic feature, it means that the semantic information of the clustering center of the above-mentioned selected alternative outline group is closest to the description information of the text to be generated. Determining the outline for the text to be generated from the selected alternative outline group can make the semantic information expressed by the determined outline for the text to be generated closest to the description information of the text to be generated, thereby improving the accuracy of determining the outline for the text to be generated.
[0147] In an embodiment of the present invention, the above S202 of selecting the outline for the text to be generated from the outlines within the alternative outline group according to the similarity between the semantic features of each outline in the alternative outline group and the first semantic feature can be implemented according to the following steps B1 - B2.
[0148] Step B1: Calculate the similarity between the semantic features of each outline in the alternative outline group and the first semantic feature.
[0149] Specifically, when calculating the similarity between the semantic features of each outline in the alternative outline group and the first semantic feature, the distance between the semantic features of each outline and the first semantic feature can be calculated, and the similarity can be determined based on the calculated distance.
[0150] Step B2: Select the first preset number of outlines from each outline group in the order of the calculated similarity of each outline from high to low as the outline for the text to be generated.
[0151] The above first preset number can be set by the staff according to experience. For example: the above first preset number can be 6, 8, etc. Taking the first preset number as 6 as an example, 6 outlines with the highest similarity can be selected.
[0152] After obtaining the similarity between the semantic features of each outline and the first semantic feature, select the first preset number of outlines in the order of the calculated similarity of each outline from high to low. It can be understood that the selected outlines can be outlines from different alternative outline groups or outlines from the same alternative outline group.
[0153] For example: Assume that the order of the similarity of each outline from high to low is: Outline 1, Outline 2, Outline 3, Outline 4, Outline 5, and the first preset quantity is 3. Then the first preset quantity of outlines selected are: Outline 1, Outline 2, Outline 3.
[0154] In this way, since the first preset quantity of outlines are selected as the outlines of the text to be generated in the order of the similarity of each outline from high to low, and the above similarity can reflect the similarity between the semantic information expressed by each outline and the description information of the text to be generated, the accuracy of selecting the outlines of the text to be generated is improved.
[0155] In an embodiment of the present invention, the following steps C1 - C4 can be followed to implement the similarity between the semantic feature of the clustering center of each outline group and the first semantic feature in S201, and select the outline group to which the outline of the text to be generated belongs from each outline group.
[0156] Step C1: Calculate the similarity between the semantic feature of the clustering center of each outline group and the first semantic feature.
[0157] Specifically, when calculating the similarity between the semantic feature of the clustering center of the outline group and the first semantic feature, the distance between the semantic feature of each clustering center and the first semantic feature can be calculated, and the similarity can be determined based on the calculated distance.
[0158] Step C2: Select the first second preset quantity of outline groups from each outline group in the order of the similarity of each calculated outline group from high to low.
[0159] The above second preset quantity can be set by the staff according to experience. For example: The above second preset quantity can be 5, 10, etc.
[0160] After obtaining the similarity between the semantic feature of the clustering center of each outline group and the first semantic feature, select the first second preset quantity of outline groups from each outline group in the order of the similarity of each calculated outline group from high to low.
[0161] For example: Assume that the order of the similarity of each outline group from high to low is: Outline Group 1, Outline Group 2, Outline Group 3, Outline Group 4, Outline Group 5, and the second preset quantity is 3. Then the first second preset quantity of outline groups selected are: Outline Group 1, Outline Group 2, Outline Group 3.
[0162] Step C3: Determine the number of outlines containing description information in each selected outline group.
[0163] The description information is the description information of the text to be generated obtained in S101.
[0164] There are two cases for the outline containing description information. One case is the outline containing all the description information of the text to be generated, and the other case is the outline containing partial description information of the text to be generated.
[0165] Specifically, when determining the number of the above-mentioned outlines, it can be first determined whether the outlines in each of the above-mentioned outline groups contain description information. In one implementation manner, the description information can be extracted from each of the outlines in the above-mentioned outline group, and it can be judged whether the extracted description information is the description information of the text to be generated. If so, it is considered that the outline contains description information.
[0166] In another implementation manner, the description information of each outline can also be extracted in advance, and the extracted description information can be stored in a database. The description information of the text to be generated is matched with the description information corresponding to each outline stored in the database. If the match is successful, it can be considered that the outline contains description information.
[0167] After it is determined that the outlines in each outline group contain description information, the number of outlines containing description information in each outline group can be counted, so as to determine the number of outlines containing description information in each of the selected outline groups.
[0168] Step C4: According to the determined number of outlines, determine the outline group to which the outline of the text to be generated belongs from the selected outline groups as the alternative outline group.
[0169] In one implementation manner, the same initial weight is assigned to each outline group. Among them, the weight of the above-mentioned outline group is used to represent the probability that the outline group is selected as the alternative outline group. When the selection weight of the outline group is larger, it means that the probability that the outline group is selected as the alternative outline group is larger; when the selection weight of the outline group is smaller, it means that the probability that the outline group is selected as the alternative outline group is smaller.
[0170] Based on the determined number of outlines, adjust and update the weights of each outline group. After the weights of each outline group are updated and adjusted, a preset number of outline groups with the highest weight values can be selected as the alternative outline groups.
[0171] For example: Suppose the selected outline groups include outline group S1 and outline group S2. Among them, the number of outlines containing description information in outline group S1 is 5, the number of outlines containing description information in outline group S2 is 3, and the initial weights of outline group S1 and outline group S2 are 10, and the preset number is 1.
[0172] Adjust the weight values of each outline group based on the number of outlines containing description information in each outline group. The adjusted weight values are respectively: the weight of outline group S1 is 15, and the weight of outline group S2 is 12. Since the weight of outline group S1 is the highest, outline group S1 can be used as the alternative outline group.
[0173] In another implementation, first, the same initial weight is assigned to each outline group, the number of pieces of description information included in each outline in each outline group is determined, and based on the determined number and the number of outlines, the weight of each outline group is adjusted, and a preset number of outline groups with the highest weight values can be selected as alternative outline groups.
[0174] For example: Suppose the selected outline groups include outline group S3 and outline group S4. Among them, the number of outlines containing description information in outline group S3 is 2, the number of outlines containing description information in outline group S4 is 3, the number of pieces of description information included in each outline in outline group S3 is 2, the number of pieces of description information included in each outline in outline group S4 is 3, and the initial weights of outline group S3 and outline group S4 are 10, and the preset number is 1
[0175] Based on the number of outlines containing description information in each outline group and the number of pieces of description information included in each outline within each outline group, the weight values of each improvement group are adjusted. The weight values of each outline group can be adjusted according to the weight corresponding to the preset number of outlines and the weight corresponding to the number of pieces of description information included.
[0176] In this way, since the alternative outline groups are determined based on the determined number of outlines, and the determined number of outlines can reflect the situation of each outline in the outline group containing description information. When the determined number of outlines is larger, it can indicate that there are a larger number of outlines in the outline group containing description information, and it can be considered that the semantic information expressed by the outlines in the outline group is close to the description information of the text to be generated. Therefore, based on the determined number of outlines, the accuracy of determining the alternative outline groups can be improved.
[0177] See Figure 3 , Figure 3 , which is a schematic flowchart of the second outline determination method provided by the embodiment of the present invention. On the basis of the above embodiment, the above method further includes the following step S104.
[0178] S104: Select the paragraphs of the text to be generated from the preset paragraphs corresponding to the outlines of the text to be generated.
[0179] The preset paragraphs can be: After extracting the outlines from the text obtained in advance, the paragraphs corresponding to each outline are obtained as the preset paragraphs corresponding to the outlines.
[0180] Taking a patent text as an example, the paragraph corresponding to the outline "Abstract of the Specification" is: the content of the abstract of the specification part, and the paragraph corresponding to the outline "Claims" is: the content of the claims part.
[0181] Specifically, an electronic device can store historical texts, extract outlines from the stored texts, and obtain the paragraphs corresponding to the extracted outlines as preset paragraphs. It can also, based on an automatic crawler system, regularly and automatically crawl the texts of a specified website and save them in a database, thereby obtaining incremental texts of the specified website, and monitor the text information on the Internet in real time, explore the sources of text data, and detect the dynamic information of text data, so as to obtain a relatively large number of texts, extract the outlines from the obtained texts, and obtain the paragraphs corresponding to each outline as the preset paragraphs corresponding to the outlines.
[0182] The preset outline and the preset paragraph are in a corresponding relationship. Specifically, one preset outline can correspond to one preset paragraph or multiple preset paragraphs. Therefore, after determining the outline of the text to be generated, the preset paragraph corresponding to the outline of the text to be generated can be determined.
[0183] When selecting the paragraph of the text to be generated from the preset paragraphs corresponding to the outline of the text to be generated, in one implementation, a paragraph can be randomly selected from the preset paragraphs corresponding to the outline of the text to be generated as the paragraph of the text to be generated.
[0184] In one embodiment of the present invention, the paragraph of the text to be generated can also be selected from the preset paragraphs corresponding to the outline of the text to be generated based on the similarity between the semantic features of the preset paragraphs corresponding to the outline of the text to be generated and the first semantic feature.
[0185] The semantic features of the preset paragraphs corresponding to the outline of the text to be generated are used to reflect the semantics expressed by the preset paragraphs. The first semantic feature is the semantic feature of the description information and is used to reflect the semantics expressed by the description information.
[0186] Specifically, the above similarity can be determined by calculating the distance between the semantic features of the preset paragraphs corresponding to the outline of the text to be generated and the first semantic feature, and the similarity is determined based on the calculated distance. The above distance can be Euclidean distance, cosine distance, etc. For example: a distance-similarity conversion algorithm can be used to convert the calculated distance into similarity.
[0187] When selecting the paragraph of the text to be generated from the preset paragraphs corresponding to the outline of the text to be generated based on the similarity between the semantic features of the preset paragraphs corresponding to the outline of the text to be generated and the first semantic feature, it can be to select the preset paragraph with the highest similarity as the paragraph of the text to be generated, or it can be to select the preset paragraphs with a similarity greater than the preset similarity as the paragraphs of the text to be generated.
[0188] Since the paragraphs of the text to be generated are determined based on the similarity between the semantic features of the preset paragraphs and the first semantic features, and the above-mentioned first semantic features can reflect the semantics expressed by the description information of the text to be generated, and the semantic features of the preset paragraphs can reflect the semantics expressed by each preset paragraph, therefore, the semantic information expressed by the paragraphs determined based on the similarity between the above-mentioned semantic features can be relatively close to the semantic information expressed by the text to be generated, so that the generated text is relatively accurate.
[0189] In this way, the paragraphs of the text to be generated are selected from the preset paragraphs corresponding to the outline of the text to be generated, so that the selected paragraphs have a high degree of fit with the outline of the text to be generated, and when generating the text subsequently, the overall logic of the text is relatively strong.
[0190] See Figure 4 , Figure 4 which is a schematic flowchart of the third outline determination method provided by the embodiment of the present invention. On the basis of the above embodiment, the above method further includes the following step S105.
[0191] Step S105: Based on the outline of the text to be generated, sort the selected paragraphs of the text to be generated to generate a text including the outline of the text to be generated and the sorted paragraphs.
[0192] When sorting the selected paragraphs of the text to be generated based on the outline of the text to be generated, the arrangement order of each paragraph can be determined based on the position information of each outline of the text to be generated in the text.
[0193] For example: Suppose the outline of the text to be generated includes a beginning outline, a middle outline, and an ending outline. Among them, the selected paragraphs include Paragraph 1, Paragraph 2, and Paragraph 3, and the beginning outline corresponds to Paragraph 1, the middle outline corresponds to Paragraph 2, and the ending outline corresponds to Paragraph 3. Since the structure of the text usually consists of a beginning outline, a middle outline, and an ending outline, the arrangement order of the selected paragraphs can be determined based on the above-mentioned outline of the text to be generated, which is in turn: Paragraph 1, Paragraph 2, Paragraph 3.
[0194] When generating a text including the outline of the text to be generated and the sorted paragraphs, each outline of the text to be generated can be used as the subtitle of the corresponding sorted paragraph, and the determined subtitles and each sorted paragraph are combined to generate a text. It can also be that each outline of the text to be generated is used as the beginning sentence of the corresponding sorted paragraph, so as to generate a text.
[0195] In this way, since the selected paragraphs are sorted based on the outline of the text to be generated, a text containing the outline of the text to be generated and the sorted paragraphs is generated. Also, since the outline can reflect the structural information of the text, sorting the selected paragraphs based on the outline of the text to be generated can make the sorted paragraphs have structure, thereby improving the structure of the generated text and making the overall logic of the generated text relatively high.
[0196] In addition, since the outline of the text to be generated is determined first, and then the paragraphs of the text to be generated are selected from the preset paragraphs corresponding to the outline of the text to be generated, the number of preset paragraphs corresponding to the outline of the text to be generated is much smaller than the total number of preset paragraphs. Therefore, the efficiency of selecting the paragraphs of the text to be generated from the preset paragraphs corresponding to the outline of the text to be generated is relatively high, thereby improving the efficiency of text generation.
[0197] In the prior art, electronic devices usually generate text in units of words. Specifically, first, according to the keywords of the text to be generated obtained, the first word of the text to be generated is determined, and based on the determined first word, the words adjacent to the first word are determined. In this way, each word of the text to be generated is determined, and the text is generated based on the determined words. Using the method of generating text in units of words in the prior art results in poor structure of the generated text and weak overall logic of the text. However, in this solution, the text is not generated word by word. Sorting the selected paragraphs based on the outline of the text to be generated can make the sorted paragraphs have structure, thereby improving the structure of the generated text and making the overall logic of the generated text relatively high.
[0198] In one embodiment of the present invention, the description information of the text to be generated in the above step S101 may include at least one of the following information: user profile, keywords, entity words, key sentences, and text type.
[0199] Specifically, the user profile is used to describe the characteristic information of the user. The user profile can describe the characteristic information of the user from different dimensions. For example, the user profile can include user attribute characteristic information, user interest characteristic information, user behavior characteristic information, and user scenario characteristic information.
[0200] Among them, the user attribute characteristic information may include the user's gender, age, location, occupation, etc.; the user interest attribute characteristic information may include the types of articles the user is interested in, the types of articles to be created, writing rules, etc.; the user behavior characteristic information includes the types of articles the user has recently read and the types of articles to be created recently; the user scenario characteristic information includes the user's current writing scenario.
[0201] Specifically, when obtaining a user profile, it can be determined based on the user identifier and the pre-stored correspondence between the user identifier and the user profile. The above-mentioned pre-stored correspondence between the user identifier and the user profile can be the real-time monitoring of the user's online and offline behavior data by the electronic device, constructing a comprehensive, accurate, and multi-dimensional user profile for each user, and determining the correspondence between the user identifier and the user profile based on the constructed user profile.
[0202] Keywords can be understood as words that express the central idea or main content of the text. Specifically, when obtaining the above keywords, it can be determined based on the frequency of each word in the description text segment of the text to be generated input by the user in the text segment. For example: The TF-IDF (term frequency–inverse document frequency) algorithm can be used to extract keywords. In the TF-IDF algorithm, the following formula is mainly used to extract keywords;
[0203]
[0204] Among them, tf i,j represents the frequency of the i-th word in the above description text segment, df i represents the number of texts containing the i-th word in the preset text library, N represents the total number of texts in the preset text library, and W i,f is the importance value of the i-th word in the above description text segment. Specifically, when W i,f is higher, it means that the importance value of the i-th word in the above description text segment is higher. When W i,f is the highest, the i-th word can be considered as a keyword.
[0205] Entity words can be understood as proper nouns that appear in the text. When obtaining entity words, bidirectional LSTM and conditional random fields can be used to extract the entity words of the description text segment of the text to be generated input by the user.
[0206] Key sentences can be understood as sentences that express the central idea or main content of the text. Specifically, when obtaining key sentences, the description text segment of the text to be generated input by the user can be segmented into several sentence units, and a graph model can be established based on the context relationship between each sentence unit. Based on the established graph model, the sentence units with higher importance are determined, so as to extract the key sentences in the text segment. The above method can be implemented based on the TextRank algorithm.
[0207] When obtaining the text type, the semantic features of the description text segment of the text to be generated input by the user can be recognized, and the recognized semantic features are matched with the semantic features of the preset text types. Based on the matching results, the text type is determined.
[0208] The following specific embodiment illustrates the specific implementation process of obtaining the paragraphs of the text to be generated. Refer to Figure 5 , Figure 5 which is a flowchart of a method for obtaining paragraphs provided by an embodiment of the present invention.
[0209] In Figure 5 , first, an optimal outline group is selected from each of the clustered outline groups based on the semantic features of the text to be generated. The above-mentioned outline groups include Outline Group 1, Outline Group 2,....
[0210] Then, the outline of the text to be generated is determined from the selected optimal outline group based on the feature information of the text to be generated. Specifically, it can be determined based on the situation that each outline in the optimal outline group contains the feature information of the text to be generated. The outlines in the above-mentioned optimal outline group include Outline 1, Outline 2,....
[0211] Secondly, from the preset paragraphs corresponding to the determined outlines, the paragraphs of the text to be generated are determined based on the feature information of the text to be generated. The preset paragraphs corresponding to the above-mentioned Outline 1 include Paragraph 11, Paragraph 12,.... The preset paragraphs corresponding to the above-mentioned Outline 2 include Paragraph 21, Paragraph 22,.... Specifically, it can be determined based on the situation that each paragraph contains the feature information of the text to be generated.
[0212] In an embodiment of the present invention, the semantic features and feature information of the above-mentioned preset outlines, and the semantic features and feature information of the preset paragraphs can be determined in the following manner. Refer to Figure 6 , Figure 6 which is a flowchart of a method for obtaining text information provided by an embodiment of the present invention.
[0213] In Figure 6 , first, data acquisition is performed. Specifically, the electronic device can, based on an automatic crawler system, monitor the text information on the Internet in real time, periodically and automatically crawl the text of a specified website and save it to the message cache message queue, and use a preset data scheduling system to discover the text data source from the message cache message queue and monitor the dynamic information of the text data crawled from the Internet in real time and concurrently.
[0214] Secondly, cleaning and filtering are performed. Specifically, operations such as removing duplicates of text similarity, spelling correction, filtering sensitive topics in the text, abnormal character processing, sentiment analysis, correcting the mixed use of traditional and simplified Chinese, correcting misspellings, unifying the text format, and deleting useless information can be carried out.
[0215] Construct the document portrait again. Specifically, the outlines and paragraphs of the text after cleaning and filtering can be extracted respectively, and the feature information in the outlines and paragraphs can be obtained respectively. The above feature information may include document type, keywords, key sentences, entity words, text body, etc. When obtaining the feature information in the outlines and paragraphs, the same processing method can be used for processing.
[0216] Specifically, when obtaining the above keywords, it can be determined based on the frequency of each word in the description text segment of the text to be generated input by the user in the text segment. For example, the TF-IDF algorithm can be used to extract keywords. When obtaining the above entity words, bidirectional LSTM and conditional random fields can be used to extract the entity words of the description text segment of the text to be generated input by the user. When obtaining the above key sentences, the description text segment of the text to be generated input by the user can be segmented into several sentence units, and a graph model can be established based on the context relationship between each sentence unit. Based on the established graph model, the sentence units with higher importance can be determined, so as to extract the key sentences in the text segment. The above method can be implemented based on the TextRank algorithm. When obtaining the above text type, the semantic features of the description text segment of the text to be generated input by the user can be recognized, and based on the matching between the recognized semantic features and the semantic features of the preset text type, the text type can be determined based on the matching result.
[0217] Finally, feature extraction and data storage are performed.
[0218] Specifically, vectorized encoding technology can be used to encode the outline or paragraph, and the encoded result can be used as the semantic feature of the outline or paragraph.
[0219] And after obtaining the above text portrait and semantic features, two sets of distributed data storage methods can be used to store the above data. One set of distributed data storage systems is used to store the feature information of the above outlines and paragraphs. The above data storage system can store the original information of the text and feature information such as keywords, word segmentation, and inverted index. Another set of distributed data storage systems is used to store the semantic features of the above outlines and paragraphs for calculating the semantic feature similarity.
[0220] In one embodiment of the present invention, the similarity between the semantic feature based on the preset outline and the first semantic feature in S102 above can be realized in the following manner, and the outline of the text to be generated is selected from the preset outline as the outline of the text to be generated.
[0221] The semantic feature of the preset outline and the first semantic feature are input into a pre-trained semantic similarity calculation model to obtain the similarity between the semantic feature of the preset outline and the first semantic feature, so as to select the outline of the text to be generated from the preset outline based on the obtained similarity.
[0222] The above semantic similarity calculation model is trained on a preset neural network model with the semantic features of a large number of sample preset outlines and the semantic features of sample texts as inputs and the similarity between the semantic features of the sample preset outline and the sample text as the training benchmark, and is used to obtain the similarity between the semantic features of the preset outline and the first semantic features.
[0223] The semantic features of the above sample preset outline and the semantic features of the sample text can use the semantic features in the semantic feature vector library above Figure X as samples.
[0224] Specifically, refer to Figure 7 , Figure 7 which is a flowchart of a process for calculating semantic similarity based on a model provided by an embodiment of the present invention.
[0225] Figure 7 The calculation is based on the semantic similarity calculation model. Specifically, when calculating the similarity between the semantic features of the preset outline and the first semantic features, first, the similarities between the semantic features of the preset outline and the first semantic features are encoded respectively to obtain encoding results; then, the obtained encoding results are connected, and the difference and inner product are calculated; then, the calculation results are processed through a fully connected layer and a Softmax layer, and the similarity between the semantic features of the preset outline and the first semantic features is output.
[0226] The following uses a specific embodiment to illustrate the specific implementation process of text generation. Refer to Figure 8 , Figure 8 which is a flowchart of a text generation method provided by an embodiment of the present invention. In Figure 8 it includes a data platform construction module, a user portrait and user intention recognition module, vectorized semantic retrieval, and a text generation module.
[0227] Among them, in the data platform construction module, it includes steps of data acquisition, data cleaning, feature information, semantic feature extraction, and data storage.
[0228] In the user portrait and intention recognition module, first, the information of the text to be generated input by the user is obtained; then, the above information is filtered, for example: filtering sensitive words and stop words in the input information to remove noise information; secondly, feature information extraction is performed on the cleaned information, and the above feature information may include keywords, entity words, and text types of the text to be generated, and semantic feature extraction is performed on the above feature information.
[0229] In the vectorized semantic retrieval module, based on the extracted semantic features, similar semantic features of the outlines and paragraphs stored in the distributed vector database are recalled.
[0230] In a text generation system, first, an outline group relatively similar to the above semantic features is determined from a distributed vector database, the extracted outline groups are sorted based on the above feature information, the top pre-set number of outline groups after sorting are selected, and paragraphs of the text to be generated are determined based on the semantic features of the above semantic features and the pre-set paragraphs corresponding to the selected outline groups, so as to generate a text based on the determined outlines and paragraphs.
[0231] In one embodiment of the present invention, the pre-set paragraph corresponding to the outline in the above step S104 is a pre-determined paragraph, and specifically, the pre-set paragraph corresponding to the outline can be obtained according to the following steps D1 - D2.
[0232] Step D1: Obtain the pre-selected text corresponding to the pre-set outline.
[0233] As can be known from the description in the above step S103, the electronic device can store historical texts, extract each outline from the historical texts, or can also use an automatic crawler system to regularly crawl texts of a specified website, extract the outlines of the crawled texts, save the above outlines to the database, and store the corresponding relationship between each outline and the text. On this basis, the electronic device can pre-select the text corresponding to the pre-set outline from the texts stored in the database as the pre-selected text corresponding to the above pre-set outline.
[0234] Step D2: Extract paragraphs from each paragraph corresponding to the pre-set outline in the text as the pre-set paragraph corresponding to the outline.
[0235] In one implementation manner, paragraphs can be randomly selected from each paragraph corresponding to the pre-set outline in the text as the pre-set paragraph corresponding to the outline, where the number of the selected paragraphs can be 1 or multiple.
[0236] In one embodiment of the present invention, the feature information of each paragraph corresponding to the pre-set outline in the text can also be determined, and based on the feature information of each paragraph, alternative paragraphs are selected from each paragraph, and the alternative paragraphs are determined as the pre-set paragraph corresponding to the outline.
[0237] The above feature information of the paragraph is used to describe the basic information of the paragraph, and the above feature information of the paragraph can include: information such as the length of each sentence in the paragraph, the punctuation composition of each sentence, and the number of sentences in the paragraph.
[0238] Taking the feature information as the length of each sentence in the paragraph as an example, when determining the above feature information, the length of each paragraph in the text can be calculated.
[0239] When selecting alternative paragraphs, it can be determined whether the characteristic information of each paragraph meets a preset information screening rule. If it meets, the paragraph is used as an alternative paragraph.
[0240] The above-mentioned preset information screening rule can be set by the staff according to experience, or the staff can extract the characteristic information of high-quality paragraphs and use the screening rule containing the high-quality characteristic information as the above-mentioned preset information screening rule.
[0241] For example: The above-mentioned preset information screening rule can include: the sentence length in the paragraph is greater than 8 bytes; the punctuation marks in the sentence include commas and full stops; the number of sentences in the paragraph is greater than 5, etc.
[0242] When the characteristics of the paragraph meet the above-mentioned preset information screening rule, the paragraph can be used as an alternative paragraph.
[0243] For example: Suppose the preset information screening rule is; retain the paragraphs with sentence length greater than 8 bytes, and the paragraph contains preset punctuation marks (commas, full stops), and the number of sentences in the paragraph is greater than 5. Suppose the paragraph is: "We should study hard and make progress every day, and also respect the elderly and care for the young. When classmates encounter difficulties, we should offer a helping hand in time. Listen carefully in class and do homework carefully after class." In this paragraph, there is a sentence "We should study hard and make progress every day" with a length greater than 8 bytes, and the paragraph contains commas and full stops, and the number of sentences in the paragraph is greater than 5. That is, the above paragraph meets the preset information screening rule, so the above paragraph can be used as an alternative paragraph.
[0244] When determining the alternative paragraph as the preset paragraph corresponding to the outline, in one implementation manner, the alternative paragraph can be directly determined as the preset paragraph corresponding to the preset outline in the text. For example: Suppose the determined alternative paragraph is paragraph 1, and paragraph 1 can be directly determined as the preset paragraph corresponding to the preset outline in the text.
[0245] In one embodiment of the present invention, for each alternative paragraph, the semantic characteristics of the alternative paragraph and the semantic characteristics of each word in the alternative paragraph can also be determined, and the determined semantic characteristics and semantic characteristics of the words are input into a pre-trained paragraph quality evaluation model to obtain the quality score value of the alternative paragraph. The alternative paragraph with the quality score value greater than the preset quality score value is used as the preset paragraph corresponding to the preset outline in the text.
[0246] The above-mentioned semantic characteristics of the alternative paragraph are used to reflect the semantics expressed by the alternative paragraph, and the semantic characteristics of the above-mentioned words are used to reflect the semantics expressed by the words.
[0247] Specifically, when determining the semantic features of each word in the above alternative paragraphs, the alternative paragraphs can be segmented to obtain each word in the alternative paragraphs, and each word in the alternative paragraphs can be vectorized to obtain the semantic features of each word in the alternative paragraphs. For example, when using the jieba word segmentation tool for segmentation, the Word2Vec (Word to Vector) model can be used to vectorize each word.
[0248] When determining the semantic features of the above alternative paragraphs, the semantic information of the alternative paragraphs can be extracted, and the semantic features of the alternative paragraphs can be determined based on the extracted semantic information.
[0249] It is also possible to determine the degree of association between each word based on the semantic features of each word in the alternative paragraph, and determine the semantic features of the alternative paragraph based on the determined degree of association. Specifically, the Attention mechanism can be used. First, based on the semantic features of each word in the alternative paragraph, the degree of association between each word is determined, and weights are assigned to each word based on the determined degree of association to obtain a weight matrix; then, based on the word vectors of each word and the above weight matrix, the weights of each word are weighted and summed to obtain a weight matrix after the weighted sum of each word; finally, based on the word vectors of each word and the above weight matrix after the weighted sum of each word, the semantic features of the alternative paragraph are determined.
[0250] The above paragraph quality evaluation model is: obtained by training a preset neural network model with the semantic features of the sample paragraphs and the semantic features of each word in the sample paragraphs as the model input, the marked quality score value of the sample paragraphs as the training benchmark, and is used to obtain the quality score value of the paragraphs. The above neural network model can be TextCNN (Text Convolutional Neural Networks).
[0251] Specifically, when inputting the semantic features of the above alternative paragraphs and the semantic features of each word in the alternative paragraph into the paragraph quality evaluation model, first, the convolutional layer in the paragraph quality evaluation model performs a convolutional operation on the above semantic features and semantic features to obtain a convolutional result; the convolutional result is pooled and softmax solved to obtain the quality score value of the alternative paragraph.
[0252] Specifically, when performing softmax (logistic regression) solution, the following formula can be used for calculation:
[0253]
[0254] where p represents the serial number of the preset quality classification, and k represents the total number of preset quality classifications. Indicates that the quality of the alternative paragraph is the quality score value of the p-th preset quality classification. Indicates that the quality of the alternative paragraph is the sum of the quality score values of each preset quality classification, a p Indicates that the quality of the alternative paragraph is the normalized quality score value of the p-th preset quality classification. When the quality score value is calculated using the above formula, the calculated quality score value is within the range of [0, 1].
[0255] In an embodiment of the present invention, the semantic features and feature information of the above preset outline, and the semantic features and feature information of the preset paragraph can be determined in the following manner. See Figure 8 , Figure 8 is a flowchart of a text information acquisition method provided by an embodiment of the present invention.
[0256] In Figure 8 , data acquisition is first performed. Specifically, the electronic device can, based on an automatic crawler system, monitor the text information on the Internet in real time, periodically and automatically crawl the text of a specified website and save it to the message cache message queue, and use a preset data scheduling system to discover the text data source from the message cache message queue, and monitor the dynamic information of the text data crawled from the Internet in real time and concurrently.
[0257] Secondly, cleaning and filtering are performed. Specifically, operations such as removing duplicate text similarity, spelling correction, filtering sensitive topics in the text, abnormal character processing, sentiment analysis, correcting the mixed use of traditional and simplified Chinese, correcting misspellings, unifying the text format, and deleting useless information can be carried out.
[0258] Thirdly, the construction of the document portrait is carried out. Specifically, the outline and paragraphs of the text after cleaning and filtering can be extracted respectively, and the feature information in the outline and paragraphs can be obtained respectively. The above feature information can include document type, keywords, key sentences, entity words, text body, etc. When obtaining the feature information in the outline and paragraphs, the same processing method can be used for processing.
[0259] Specifically, when obtaining the above-mentioned keywords, it can be determined based on the frequency of each word in the description text segment of the text to be generated input by the user. For example, the TF-IDF algorithm can be used to extract keywords. When obtaining the above-mentioned entity words, bidirectional LSTM (Long Short-Term Memory) and conditional random fields can be used to extract the entity words in the description text segment of the text to be generated input by the user. When obtaining the above-mentioned key sentences, the description text segment of the text to be generated input by the user can be segmented into several sentence units, and a graph model can be established based on the context relationship between each sentence unit. Based on the established graph model, the sentence units with higher importance can be determined, so as to extract the key sentences in the text segment. The above method can be implemented based on the TextRank algorithm. When obtaining the above-mentioned text type, the semantic features of the description text segment of the text to be generated input by the user can be recognized, and based on the recognized semantic features, they can be matched with the semantic features of the preset text types, and the text type can be determined based on the matching result.
[0260] Finally, feature extraction and data storage are performed.
[0261] Specifically, vectorized encoding technology can be used to encode the outline or paragraph, and the encoded result can be used as the semantic feature of the outline or paragraph.
[0262] And after obtaining the above-mentioned text portrait and semantic features, two sets of distributed data storage methods can be used to store the above data. One set of distributed data storage systems is used to store the feature information of the above-mentioned outline and paragraph. The above data storage system can store the original information of the text and feature information such as keywords, word segmentation, and inverted index. Another set of distributed data storage systems is used to store the semantic features of the above-mentioned outline and paragraph for calculating the semantic feature similarity.
[0263] See Figure 9a , Figure 9a which is a schematic flowchart of a process for obtaining a preset paragraph provided by an embodiment of the present invention.
[0264] In Figure 9a , a large number of document data are obtained in the first step, and the above-mentioned document data includes a large amount of text.
[0265] In the second step, paragraphs in each text are extracted to form a paragraph list.
[0266] In the third and fourth steps, based on the preset paragraph description information such as length, punctuation marks, and the number of sentences, the paragraphs in the above paragraph list are roughly extracted to obtain alternative paragraphs.
[0267] In the fifth and sixth steps, the Word2Vec model is used to determine the semantic features of each word in each alternative paragraph, and the Attention mechanism is used to determine the semantic features of each alternative paragraph. Then, the determined semantic features and semantic features are input into a pre-trained TextCNN model to obtain the quality score value of the alternative paragraph.
[0268] In the seventh step, based on the quality score values of each alternative paragraph, the alternative paragraphs with quality score values greater than the preset quality threshold are used as the preset paragraphs.
[0269] In this way, since the quality score value of the paragraph is determined by combining deep learning methods, the quality of the paragraph can be accurately determined based on the above quality score value.
[0270] See Figure 9b , Figure 9b which is a schematic flowchart of a process for evaluating the quality of alternative paragraphs provided by an embodiment of the present invention.
[0271] In Figure 9b , in the order of the arrow directions, first, a list of paragraphs after rough extraction is obtained, that is, each alternative paragraph.
[0272] Secondly, each alternative paragraph is processed in a batch manner.
[0273] Then, for each alternative paragraph, each word in the alternative paragraph is obtained by using a word segmentation method, and each word is vectorized by using the Word2Vec method to obtain the word vectors of each word in the alternative paragraph. Furthermore, based on the obtained word vectors and the Attention mechanism, an attention feature map (Attention FeatureMap) of the alternative paragraph is obtained.
[0274] Finally, the word vectors of each word in the above alternative paragraph and the attention feature map are input into a CNN (Convolutional Neural Networks) model. After convolution operations, pooling, and softmax in the CNN model, the quality score value of the alternative paragraph is output.
[0275] In this way, since each alternative paragraph is processed in a batch manner, the processing efficiency is improved, and further the efficiency of paragraph acquisition is improved.
[0276] See Figure 9c , Figure 9c which is a schematic flowchart of an Attention mechanism provided by an embodiment of the present invention. In Figure 9cAmong them, first, a dot product operation (dot) is performed based on the semantic features of each word in the alternative paragraphs to obtain the word weights (weight) of each word in the alternative paragraphs; the softmax solution is performed on the word weights of each word in the alternative paragraphs to obtain the weight normalization result; then, the semantic features of each word in the alternative paragraphs are weighted and summed (summarize) with the weight normalization result to output a feature map (feature map); based on the above feature map and the semantic features of each word in the alternative paragraphs, a final feature map (final feature map) is obtained.
[0277] Corresponding to the above outline determination method, an embodiment of the present invention further provides an outline determination device.
[0278] See Figure 10 , Figure 10 FIG. is a schematic structural diagram of the first outline determination device provided by an embodiment of the present invention, and the above device includes the following modules 1001-1003.
[0279] An information acquisition module 1001, configured to acquire description information of a text to be generated;
[0280] A feature acquisition module 1002, configured to acquire semantic features of the description information as first semantic features;
[0281] An outline selection module 1003, configured to select an outline of the text to be generated from the preset outlines based on the semantic features of the preset outlines and the first semantic features.
[0282] In an embodiment of the present invention, the above outline selection module 1003 is specifically configured to select an outline of the text to be generated from the preset outlines based on the similarity between the semantic features of the preset outlines and the first semantic features.
[0283] See Figure 11 , Figure 11 FIG. is a schematic structural diagram of an outline selection module 1003 provided by an embodiment of the present invention, and the above module includes the following sub-modules 10031-10032.
[0284] An outline group selection sub-module 10031, configured to select, from each outline group, an outline group to which the outline of the text to be generated belongs as an alternative outline group based on the similarity between the semantic features of the clustering centers of each outline group and the first semantic features, where each outline group is an outline group obtained by clustering according to the similarity between the semantic features of the outlines;
[0285] An outline selection sub-module 10032, configured to select an outline of the text to be generated from each outline in the alternative outline group according to the similarity between the semantic features of each outline in the alternative outline group and the first semantic feature.
[0286] In an embodiment of the present invention, the above-mentioned outline selection sub-module 10032 is specifically configured to select, from the alternative outline group, an alternative outline group with the highest similarity between the semantic feature of the clustering center and the first semantic feature; and select an outline of the text to be generated from each outline in the selected alternative outline group according to the similarity between the semantic features of each outline in the selected alternative outline group and the first semantic feature.
[0287] In an embodiment of the present invention, the above-mentioned outline selection sub-module 10032 is specifically configured to calculate the similarity between the semantic features of each outline in the alternative outline group and the first semantic feature; and select the first preset number of outlines from each outline in descending order of the calculated similarity corresponding to each outline as the outline of the text to be generated.
[0288] In an embodiment of the present invention, the above-mentioned outline group selection sub-module 10031 includes:
[0289] A similarity calculation unit, configured to calculate the similarity between the semantic feature of the clustering center of each outline group and the first semantic feature;
[0290] An outline group selection unit, configured to select the first second preset number of outline groups from each outline group in descending order of the calculated similarity corresponding to each outline group;
[0291] A quantity determination unit, configured to determine the number of outlines of the outlines including the description information in each selected outline group;
[0292] An outline group determination unit, configured to determine the outline group to which the outline of the text to be generated belongs from the selected outline groups according to the determined number of outlines.
[0293] See Figure 12 , Figure 12 which is a schematic structural diagram of a second outline determination device provided by an embodiment of the present invention. The above-mentioned device further includes a paragraph selection module 1004.
[0294] The above-mentioned paragraph selection module 1004 is specifically configured to select a paragraph of the text to be generated from a preset paragraph corresponding to the outline of the text to be generated.
[0295] In one embodiment of the present invention, the above-mentioned paragraph selection module 1004 is specifically configured to select the paragraphs of the text to be generated from the preset paragraphs corresponding to the outline of the text to be generated based on the similarity between the semantic features of the preset paragraphs corresponding to the outline of the text to be generated and the first semantic feature.
[0296] In one embodiment of the present invention, the above-mentioned device further includes a preset paragraph determination module, and the preset paragraph determination module includes:
[0297] A text acquisition sub-module, configured to acquire pre-selected text corresponding to a preset outline;
[0298] A paragraph determination sub-module, configured to extract paragraphs from each paragraph corresponding to the preset outline in the text as the preset paragraphs corresponding to the outline.
[0299] In one embodiment of the present invention, the above-mentioned paragraph determination sub-module includes:
[0300] An information determination unit, configured to determine the feature information of each paragraph corresponding to the preset outline in the text;
[0301] A paragraph determination unit, configured to select alternative paragraphs from the respective paragraphs based on the feature information of the respective paragraphs, and determine the alternative paragraphs as the preset paragraphs corresponding to the preset outline in the text.
[0302] In one embodiment of the present invention, the above-mentioned paragraph determination unit is specifically configured to, for each alternative paragraph, determine the semantic feature of the alternative paragraph and the semantic feature of each word in the alternative paragraph, and input the determined semantic feature and semantic feature into a pre-trained paragraph quality evaluation model to obtain a quality score value of the alternative paragraph, and use the paragraph with a quality score value greater than a preset quality score threshold as the preset paragraph corresponding to the preset outline in the text;
[0303] Wherein, the paragraph quality evaluation model is: a neural network model preset trained with the semantic features of sample paragraphs and the semantic features of each word in the sample paragraphs as model inputs and the marked quality score values of the sample paragraphs as training benchmarks to obtain the quality score values of paragraphs.
[0304] See Figure 13 , Figure 13 FIG. 3 is a schematic structural diagram of a third outline determination device provided by an embodiment of the present invention. The above-mentioned device further includes a text generation module 1005.
[0305] The above-mentioned text generation module 1005 is specifically configured to sort the selected paragraphs of the text to be generated based on the outline of the text to be generated, and generate a text including the outline of the text to be generated and the sorted paragraphs.
[0306] In one embodiment of the present invention, the above-described description information includes at least one of the following information: user profile, keyword, entity word, key sentence, and text type.
[0307] Corresponding to the above outline determination method, an embodiment of the present invention further provides an electronic device.
[0308] See Figure 14 , Figure 14 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention, including a processor 1401, a communication interface 1402, a memory 1403, and a communication bus 1404. Among them, the processor 1401, the communication interface 1402, and the memory 1403 complete mutual communication through the communication bus 1404.
[0309] The memory 1403 is used to store a computer program.
[0310] When the processor 1401 is used to execute the program stored on the memory 1403, it implements the outline determination method provided by the embodiment of the present invention.
[0311] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0312] The communication interface is used for communication between the above electronic device and other devices.
[0313] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0314] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0315] In another embodiment provided by the present invention, there is also provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the outline determination method provided by the embodiment of the present invention is implemented.
[0316] In another embodiment provided by the present invention, there is also provided a computer program product containing instructions, and when it runs on a computer, it causes the computer to implement the outline determination method provided by the embodiment of the present invention when executed.
[0317] In the above embodiment, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a Solid State Disk (SSD)).
[0318] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.
[0319] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the embodiments of the apparatus, electronic device, and computer-readable storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.
[0320] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.
Claims
1. A method for determining an outline, characterized in that The method includes: Obtaining the description information of the text to be generated; Obtaining the semantic features of the description information as the first semantic features; Calculating the similarity between the semantic features of the clustering centers of each outline group and the first semantic features; selecting the first second preset number of outline groups from each outline group in the order of the calculated similarities of each outline group from high to low; determining the number of outlines of the outlines containing the description information in each selected outline group; determining the outline group to which the outline of the text to be generated belongs from the selected outline groups according to the determined number of outlines as the alternative outline group, where each of the outline groups is an outline group obtained by clustering according to the similarity between the semantic features of the outlines; Selecting the outline of the text to be generated from each outline in the alternative outline group according to the similarity between the semantic features of each outline in the alternative outline group and the first semantic features.
2. The method according to claim 1, wherein The selecting the outline of the text to be generated from each outline in the alternative outline group according to the similarity between the semantic features of each outline in the alternative outline group and the first semantic features includes: Selecting from the alternative outline groups the alternative outline group with the highest similarity between the semantic features of the clustering center and the first semantic features; Selecting the outline of the text to be generated from each outline in the selected alternative outline group according to the similarity between the semantic features of each outline in the selected alternative outline group and the first semantic features.
3. The method according to claim 1, wherein The selecting the outline of the text to be generated from each outline in the alternative outline group according to the similarity between the semantic features of each outline in the alternative outline group and the first semantic features includes: Calculating the similarity between the semantic features of each outline in the alternative outline group and the first semantic features; Selecting the first first preset number of outlines from each outline in the order of the calculated similarities of each outline from high to low as the outline of the text to be generated.
4. The method according to claim 1, characterized in that The method further includes: Selecting the paragraphs of the text to be generated from the preset paragraphs corresponding to the outline of the text to be generated.
5. The method according to claim 4, characterized in that The selecting the paragraphs of the text to be generated from the preset paragraphs corresponding to the outline of the text to be generated includes: Selecting the paragraphs of the text to be generated from the preset paragraphs corresponding to the outline of the text to be generated based on the similarity between the semantic features of the preset paragraphs corresponding to the outline of the text to be generated and the first semantic features.
6. The method according to claim 5, characterized in that, The preset paragraphs are pre-determined paragraphs, including the following steps: Obtaining the pre-selected text corresponding to the preset outline; Extracting paragraphs from each paragraph corresponding to the preset outline in the text as the preset paragraphs corresponding to the outline.
7. The method according to claim 6, wherein The extracting paragraphs from each paragraph corresponding to the preset outline in the text as the preset paragraphs corresponding to the outline includes: Determining the feature information of each paragraph corresponding to the preset outline in the text; Selecting alternative paragraphs from each of the paragraphs based on the feature information of each paragraph, and determining the alternative paragraphs as the preset paragraphs corresponding to the preset outline in the text.
8. The method according to claim 7, wherein Said determining the alternative paragraph as the preset paragraph corresponding to the preset outline in the said text includes: For each alternative paragraph, determining the semantic feature of the alternative paragraph and the semantic feature of each word in the alternative paragraph; Inputting the determined semantic feature and semantic feature of each word of each alternative paragraph into a pre-trained paragraph quality evaluation model, obtaining the quality score value of each alternative paragraph, and taking the paragraph with the quality score value greater than the preset quality score threshold as the preset paragraph corresponding to the preset outline in the said text; Wherein, the paragraph quality evaluation model is: a neural network model preset trained with the semantic feature of the sample paragraph and the semantic feature of each word in the sample paragraph as the model input, and the marked quality score value of the sample paragraph as the training benchmark, and is used to obtain the quality score value of the paragraph.
9. The method according to any one of claims 4-8, characterized in that, The said method further includes: Based on the outline of the to-be-generated text, sorting the selected paragraphs of the to-be-generated text, and generating a text including the outline of the to-be-generated text and the sorted paragraphs.
10. The method according to any one of claims 4-8, characterized in that, The said description information includes at least one of the following information: user portrait, keyword, entity word, key sentence, and text type.
11. An outline determination device, characterized in that, The said device includes: An information obtaining module, configured to obtain the description information of the to-be-generated text; A feature obtaining module, configured to obtain the semantic feature of the said description information as the first semantic feature; An outline selection module, configured to select the outline of the to-be-generated text from the preset outlines based on the similarity between the semantic feature of the preset outline and the first semantic feature; The said outline selection module includes: An outline group selection sub-module, configured to select the outline group to which the outline of the to-be-generated text belongs from each outline group as the alternative outline group based on the similarity between the semantic feature of the clustering center of each outline group and the first semantic feature, wherein, each of the said outline groups is: an outline group clustered according to the similarity between the semantic features of the outlines; An outline selection sub-module, configured to select the outline of the to-be-generated text from each outline of the alternative outline group according to the similarity between the semantic feature of each outline in the alternative outline group and the first semantic feature; The said outline group selection sub-module includes: A similarity calculation unit, configured to calculate the similarity between the semantic feature of the clustering center of each outline group and the first semantic feature; An outline group selection unit, configured to select the first second preset number of outline groups from each outline group in the order of the calculated similarity corresponding to each outline group from high to low; A quantity determination unit, configured to determine the number of outlines of the outline containing the said description information in each selected outline group; An outline group determination unit, configured to determine the outline group to which the outline of the to-be-generated text belongs from the selected outline groups according to the determined number of outlines.
12. An electronic device, characterized in that, Including a processor, a communication interface, a memory, and a communication bus, wherein, the processor, the communication interface, and the memory complete communication with each other through the communication bus; The memory is used for storing a computer program; The processor, when executing the program stored on the memory, implements the method steps described in any one of claims 1-10.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method steps described in any one of claims 1-10.
Citation Information
Patent Citations
Method and device used for generating article
CN106970898A
Military official document automatic generation system and method
CN112148857A