Artificial Intelligence-Based Official Document Content Generation Method, Device, Terminal Device, and Storage Medium
By performing sentence recognition, keyword replacement and text clustering on official document content, and combining rules to extract the target official document content, the problem of irregular format and language in official document generation is solved, and the quality and work efficiency of official document are improved.
Patent Information
- Application Number
- CN202510502214.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-22
AI Technical Summary
When generating official documents content, it is difficult for the prior art to ensure format specifications and language rigor, resulting in the low quality of the generated official documents and cannot meet actual work needs.
By performing statement recognition and keyword replacement on the first text, text clustering and similar text analysis are carried out, and target official document content is generated based on rules to ensure the standardization of the format and language.
The content of the generated official documents is more in line with the norms and rigor, which improves the quality of official documents, meets actual work needs, and improves work efficiency.
Smart Images

Figure CN120012729B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, device, terminal device, and storage medium for generating official document content based on artificial intelligence. Background Art
[0002] In actual application scenarios, taking government agencies and enterprises and institutions as examples, official documents are extremely important tools for transmitting information up and down and recording matters. Official documents carry key functions such as information transmission and decision-making deployment, and have strict format specifications and rigorous language expression requirements. However, if a text generation model is used to generate official document content, the generated official document may deviate from the specification in terms of format due to the randomness of the text generation model, and it is difficult to meet the standards in aspects such as page settings and layout; it is also difficult to conform to the rigor of official documents in terms of language expression, and problems such as inappropriate word usage and casual expression may occur. Therefore, the content generated by the existing technology when generating official document content is inconsistent with the actual needs, resulting in low-quality official documents and difficulty in meeting the needs of assisting actual work. Summary of the Invention
[0003] The main purpose of the embodiments of the present invention is to provide a method, device, terminal device, and storage medium for generating official document content based on artificial intelligence, aiming to solve the problem that the quality of official documents generated in the related technology when generating official document content is low and it is difficult to meet the needs of assisting actual work.
[0004] In a first aspect, the embodiments of the present invention provide a method for generating official document content based on artificial intelligence, including:
[0005] Performing sentence recognition on a first text to obtain key sentences and the first positions of the key sentences;
[0006] Performing keyword recognition on the key sentences to obtain first words and the first types of the first words;
[0007] Performing word replacement on the first words in the first text according to the first type to obtain a second text;
[0008] Performing text clustering on the second text to obtain a target clustering result, and performing similar text analysis on each sub-cluster in the target clustering result to obtain the general sentences of the sub-cluster and the second positions of the general sentences;
[0009] Performing keyword recognition on the general sentences to obtain second words and the second types of the second words;
[0010] Extracting a first sentence rule from the key sentences according to the first words and the first type;
[0011] Extract the second statement rule by extracting rules from the general statement according to the second word and the second type;
[0012] Obtain the target keyword of the event to be announced and the third type of the target keyword;
[0013] Generate the target official document content of the event to be announced according to the target keyword and the third type, in combination with the first statement rule and the second statement rule, at the first position and the second position.
[0014] In a second aspect, an embodiment of the present invention provides an official document content generation device based on artificial intelligence, including:
[0015] A statement recognition module, configured to perform statement recognition on a first text to obtain a key sentence and the first position of the key sentence;
[0016] A first word recognition module, configured to perform keyword recognition on the key sentence to obtain a first word and the first type of the first word;
[0017] A replacement processing module, configured to perform word replacement on the first word in the first text according to the first type to obtain a second text;
[0018] A clustering analysis module, configured to perform text clustering on the second text to obtain a target clustering result, and perform similar text analysis on each sub-cluster in the target clustering result to obtain the general statement of the sub-cluster and the second position of the general statement;
[0019] A second word recognition module, configured to perform keyword recognition on the general statement to obtain a second word and the second type of the second word;
[0020] A first rule extraction module, configured to extract rules from the key sentence according to the first word and the first type to obtain a first statement rule;
[0021] A second rule extraction module, configured to extract rules from the general statement according to the second word and the second type to obtain a second statement rule;
[0022] A data acquisition module, configured to obtain the target keyword of the event to be announced and the third type of the target keyword;
[0023] An official document generation module, configured to generate the target official document content of the event to be announced according to the target keyword and the third type, in combination with the first statement rule and the second statement rule, at the first position and the second position.
[0024] In a third aspect, an embodiment of the present invention further provides a terminal device, which includes a processor, a memory, a computer program stored on the memory and executable by the processor, and a data bus for realizing connection communication between the processor and the memory. When the computer program is executed by the processor, the steps of any one of the official document content generation methods based on artificial intelligence provided in the specification of the present invention are realized.
[0025] In a fourth aspect, an embodiment of the present invention further provides a storage medium for computer-readable storage, characterized in that the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to realize the steps of any one of the official document content generation methods based on artificial intelligence provided in the specification of the present invention.
[0026] An embodiment of the present invention provides an official document content generation method, device, terminal device, and storage medium based on artificial intelligence. The method includes: performing sentence recognition on a first text to obtain a key sentence and a first position of the key sentence, so as to quickly locate the core content corresponding to the first text and provide good support for subsequent rule extraction; performing keyword recognition on the key sentence to obtain a first word and a first type of the first word; performing word replacement on the first word in the first text according to the first type to obtain a second text, so as to realize different expressions corresponding to the same type under different words, and further reduce the interference of different words and provide good support for subsequent text clustering. Then, performing text clustering on the second text to obtain a target clustering result, and performing similar text analysis on each sub-cluster in the target clustering result to obtain a general sentence of the sub-cluster and a second position of the general sentence. Then, performing keyword recognition on the general sentence to obtain a second word and a second type of the second word, so as to extract a first sentence rule from the key sentence according to the first word and the first type; and extracting a second sentence rule from the general sentence according to the second word and the second type; obtaining a target keyword of the event to be announced and a third type of the target keyword; generating a target official document content of the event to be announced according to the target keyword and the third type in combination with the first sentence rule and the second sentence rule according to the first position and the second position, so as to generate a relatively standardized and consistent official document content, further improve the quality of the generated official document, provide support for further improving the work efficiency of relevant personnel, and also solve the problem that the quality of the official document generated in the related technology when generating the official document content is low and it is difficult to meet the actual work assistance requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0028] Figure 1 It is a schematic flowchart of a method for generating official document content based on artificial intelligence provided by an embodiment of the present invention;
[0029] Figure 2 It is a schematic block diagram of the module structure of another device for generating official document content based on artificial intelligence provided by an embodiment of the present invention;
[0030] Figure 3 It is a schematic block diagram of the structure of a terminal device provided by an embodiment of the present invention. Specific Embodiments
[0031] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0032] The flowchart shown in the drawings is only an example, and does not necessarily include all contents and operations / steps, nor does it necessarily execute in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged. Therefore, the actual execution order may be changed according to the actual situation.
[0033] It should be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0034] Embodiments of the present invention provide a method, a device, a terminal device, and a storage medium for generating official document content based on artificial intelligence. Among them, the method for generating official document content based on artificial intelligence can be applied to a terminal device, and the terminal device can be an electronic device such as a tablet computer, a notebook computer, a desktop computer, a personal digital assistant, and a wearable device. The terminal device can be a server or a server cluster.
[0035] The following will describe in detail some embodiments of the present invention with reference to the drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0036] Please refer to Figure 1 , Figure 1 , which is a schematic flow chart of a method for generating official document content based on artificial intelligence provided by an embodiment of the present invention.
[0037] As Figure 1 shown, the method for generating official document content based on artificial intelligence includes steps S101 to S109.
[0038] Step S101: Perform sentence recognition on the first text to obtain key sentences and the first positions of the key sentences.
[0039] Exemplarily, the first text is a historical official document text obtained from a database that has been published or announced. Then, a target title is obtained from the historical official document text, and the historical official document text is segmented into multiple segmented sentences. Then, the degree of association or similarity between each segmented sentence and the target title is judged according to the cosine similarity, so as to determine the segmented sentence corresponding to the maximum degree of association or similarity as the corresponding key sentence.
[0040] Exemplarily, when obtaining the sentence numbers corresponding to each segmented sentence after segmenting the historical official document text, and then determining the segmented sentence corresponding to the maximum degree of association or similarity as the key sentence, the sentence number corresponding to the segmented sentence is determined as the first position corresponding to the key sentence.
[0041] In some embodiments, the performing sentence recognition on the first text to obtain key sentences and the first positions of the key sentences includes: extracting keywords from the first text to obtain the third words corresponding to the first text and the word weights corresponding to the third words; segmenting the first text into initial sentences, and determining the sentence weights corresponding to the initial sentences according to the third words and the word weights; determining any one sentence in the initial sentences as the first sentence, and obtaining the second sentence corresponding to the adjacent state of the first sentence from the initial sentences; obtaining the first weight corresponding to the first sentence and the second weight corresponding to the second sentence from the sentence weights; adjusting the first weight according to the second weight to obtain the target weight corresponding to the first sentence; performing sentence recognition on the initial sentences according to the target weight to obtain the key sentences, and obtaining the first positions corresponding to the key sentences from the first text according to the key sentences.
[0042] Exemplarily, use the term frequency-inverse document frequency algorithm to extract keywords from the first text to obtain the third words corresponding to the first text, and obtain the word frequency corresponding to the third words from the first text. Then, determine the word weights corresponding to the third words according to the word frequency. Among them, the word weights can also be further adjusted according to the positions where the third words appear.
[0043] Exemplarily, the first text is segmented into a plurality of initial statements according to sentence delimiters such as full stops, question marks, exclamation marks, etc. Thus, for each initial statement, the word weights of the third words included in the initial statement are accumulated to obtain the statement weight of the initial statement. For example, if an initial statement contains three keywords with word weights of 0.3, 0.2, and 0.1, then the statement weight of this initial statement is 0.3 + 0.2 + 0.1 = 0.6.
[0044] Exemplarily, one sentence is sequentially selected from the plurality of initial statements in order as the first statement. For example, the first initial statement is first selected as the first statement. If the first statement is the first statement in the first text, then its adjacent second statement is the statement immediately following it; if the first statement is not the first statement, then its adjacent second statement can be the statement before it or the statement after it, which is specifically determined according to actual requirements.
[0045] Exemplarily, after determining the first statement and the second statement, the first weight corresponding to the first statement and the second weight corresponding to the second statement are obtained from the statement weights, and thus a ratio coefficient is calculated according to the second weight and the first weight. When the ratio coefficient is greater than or equal to a preset value, it indicates that the gap between the first weight and the second weight is large, and the first weight can be directly determined as the target weight without processing the first weight according to the second weight, that is, the importance of the first statement and the second statement has been significantly distinguished. If additional weight adjustment is performed, the original reasonable weight relationship may be damaged, resulting in deviation in key sentence recognition. When the ratio coefficient is less than the preset value, it indicates that the gap between the first weight and the second weight is small. First, the difference value between the first weight and the second weight is calculated, and then an adjustment coefficient is determined to combine the difference value to downwardly adjust the smaller data of the first weight and the second weight, and upwardly adjust the larger data of the first weight and the second weight, that is, when the gap between the first weight and the second weight is small, the smaller weight is downwardly adjusted and the larger weight is upwardly adjusted. This can further widen the gap between the two, making the weight of important statements more prominent, so that in the subsequent key sentence recognition process, it is easier to screen out truly important statements, and thus obtain the target weight corresponding to the adjusted first statement.
[0046] Exemplarily, all the initial statements are respectively used as the first statement for weight adjustment to obtain the target weight corresponding to each initial statement, and a screening threshold is determined, and then the initial statements with target weights greater than the threshold are determined as key sentences.
[0047] Exemplarily, obtain the statement numbers corresponding to each initial statement after sentence segmentation of the first text, and then obtain the statement numbers corresponding to the key sentences from the statement numbers, so as to determine the statement numbers as the first positions corresponding to the key sentences.
[0048] In some embodiments, the obtaining the target weight corresponding to the first statement by adjusting the weight of the first weight according to the second weight includes: obtaining a first sequence of statements corresponding to the positions to the left of the second statement from the initial statements, and obtaining a second sequence of statements corresponding to the positions to the right of the second statement from the initial statements; obtaining the associated keywords corresponding to each first sub-statement in the first sequence of statements from the third words, and obtaining the relevant weights corresponding to the associated keywords from the word weights; obtaining the third weights corresponding to each first sub-statement in the first sequence of statements from the statement weights and obtaining the fourth weights corresponding to each second sub-statement in the second sequence of statements from the statement weights; adjusting the weight of the first weight according to the relevant weights in combination with the third weights and the fourth weights by using the second weight to obtain the target weight corresponding to the first statement; wherein, the target weight is obtained according to the following formula:
[0049]
[0050] Wherein, represents the target weight corresponding to the i-th first statement, represents the adjustment parameter, represents the first weight corresponding to the i-th first statement, n represents the number of sentences corresponding to the first sequence of statements, y represents the number of associated keywords corresponding to the h-th first sub-statement, represents the relevant weight corresponding to the k-th associated keyword corresponding to the h-th first sub-statement, represents the third weight corresponding to the h-th first sub-statement, g represents the number of sentences corresponding to the second sequence of statements, represents the fourth weight corresponding to the t-th second sub-statement, represents the second weight of the second statement corresponding to the adjacent state of the i-th first statement.
[0051] Exemplarily, after performing data segmentation on the first text to obtain a plurality of initial statements, store the plurality of initial statements in the order in the first text to obtain a data storage result. Thus, after determining the second statement corresponding to the first statement, find the position of the second statement from the data storage result, and then extract all the statements to the left of the second statement to form a first sequence of statements; extract all the statements to the right of the second statement to form a second sequence of statements.
[0052] Exemplarily, according to the third word, associated keywords associated with each first sub - statement in the first sequence of statements are found through word matching. Then, according to the corresponding relationship between the third word and the word weights, the relevant weights corresponding to these associated keywords are searched for.
[0053] Exemplarily, according to the corresponding relationship between the initial statement and the statement weight, the third weight corresponding to each first sub - statement in the first sequence of statements and the fourth weight corresponding to each second sub - statement in the second sequence of statements are respectively obtained from the statement weights.
[0054] Exemplarily, using the following formula, the first weight is adjusted by the second weight by combining the relevant weight, the third weight, and the fourth weight to obtain the target weight corresponding to the first statement;
[0055]
[0056] where, represents the target weight corresponding to the i - th first statement, represents the adjustment parameter, represents the first weight corresponding to the i - th first statement, n represents the number of sentences in the first sequence of statements, y represents the number of associated keywords corresponding to the h - th first sub - statement, represents the relevant weight corresponding to the k - th associated keyword corresponding to the h - th first sub - statement, represents the third weight corresponding to the h - th first sub - statement, g represents the number of sentences in the second sequence of statements, represents the fourth weight corresponding to the t - th second sub - statement, represents the second weight of the second statement corresponding to the adjacent state of the i - th first statement.
[0057] Exemplarily, the adjustment parameter is a preset value used to control the degree of adjustment; thus, through the weight information of the left - hand statement sequence and the right - hand statement sequence of the second statement and the relevant weights corresponding to the associated keywords, the complete information corresponding to the first statement and the first text is fused with the help of the second weight corresponding to the second statement, so as to more accurately obtain the target weight corresponding to the first statement, and further provide better support for obtaining the key sentence subsequently.
[0058] Specifically, when calculating the target weight corresponding to the first statement, the complete information corresponding to the first statement and the first text is fused with the help of the second statement, thereby making the weight adjustment more comprehensive and flexible, and further improving the reliability and accuracy of the weight calculation.
[0059] Step S102: Identify keywords for the key sentence to obtain the first word and the first type of the first word.
[0060] Exemplarily, a named entity recognition model is used to identify key sentences, obtaining the word type identifier corresponding to each word in the key sentence, and determining the first word corresponding to the key sentence and the first type corresponding to the first word according to the word type identifier.
[0061] For example, the named entity recognition model is a bidirectional long short-term memory network and conditional random field model based on deep learning. Then, a pre-trained model is fine-tuned using a labeled dataset, where the dataset contains a large number of sentences, and each word in the sentence is labeled with the corresponding word type, such as person name, place name, organization name, date, others, etc. During the training process, the parameters of the model are continuously adjusted to improve the recognition accuracy of the model for word types. Then, the key sentence is input into the trained named entity recognition model, and the model will analyze each word and output the corresponding word type identifier. Then, the words corresponding to the word type identifiers of person name, place name, organization name, and date are determined as the first words, and the word type identifier corresponding to the first word is determined as the first type.
[0062] Step S103: Perform word replacement on the first word in the first text according to the first type to obtain a second text.
[0063] Exemplarily, the first types include person name, place name, organization name, time, etc. Then, a common target replacement word is defined for each first type. For example, for the person name type, it can be uniformly replaced with "someone"; for the place name type, it is replaced with "somewhere". Thus, the defined different first types and their corresponding target replacement words are organized into a replacement word library for convenient subsequent search and use.
[0064] Exemplarily, the previously mentioned named entity recognition model is used to process the first text, identifying the first word and the first type corresponding to the first word. Then, the corresponding target replacement word is searched in the replacement word library according to the first type, and then the first word in the first text is replaced with the target replacement word to obtain the replaced second text. After replacing the first word corresponding to the first type with a common expression, the second text can be applied to a wider range of scenarios, providing good support for subsequent text clustering.
[0065] For example, for the first text "Zhang San went to Beijing yesterday", through named entity recognition, it can be determined that "Zhang San" is a person name and "Beijing" is a place name. Then, the corresponding target replacement words are searched in the replacement word library. The target replacement word corresponding to the person name is "someone", and the target replacement word corresponding to the place name is "somewhere". Then, the replaced second text can be obtained as "someone went to somewhere yesterday".
[0066] Step S104: Perform text clustering based on the second text to obtain a target clustering result, and perform similar text analysis on each sub-cluster in the target clustering result to obtain the general statement of the sub-cluster and the second position of the general statement.
[0067] Exemplarily, using word embeddings such as Word2Vec, GloVe, etc., the corresponding words in the second text are mapped to a low-dimensional vector space to obtain the word embedding vectors corresponding to each word in the second text. Then, by calculating the average or weighted average of the word embedding vectors of the words in the second text to represent the second text. After converting each second text into a feature vector, these vectors are combined into a feature matrix. Each row of the matrix represents a text, and each column represents a feature. Thus, using a clustering algorithm such as hierarchical clustering to perform text clustering based on the feature matrix to obtain the target clustering result.
[0068] Exemplarily, since different types of words have been unified using general words in the second text, the key words are already unified when performing clustering based on the second text to obtain the target clustering result. Based on this, the texts included in each sub-cluster in the target clustering result often have the same or similar meanings. In view of this, when performing similar text analysis on each sub-cluster in the target clustering result, the general statement corresponding to each sub-cluster can be summarized.
[0069] Exemplarily, calculate the similarity of the statements between any two sub-texts in the sub-texts corresponding to each sub-cluster in the target clustering result to obtain the closest statement between these two sub-texts. Then, obtain all the closest statements between any two sub-texts in the sub-cluster and perform intersection processing on all the closest statements. Then, determine the statement with the largest number in the intersection processing result as the general statement corresponding to the sub-cluster.
[0070] For example, for any two sub-texts in each sub-class cluster, the similarity between each statement of one sub-text and all statements of the other sub-text is calculated. For example, if sub-text A has 3 statements and sub-text B has 4 statements, then 3×4 = 12 similarity calculations are required. Thus, among the results of each group of statement similarity calculations, the statement pair with the highest similarity is found, and these statement pairs are the closest statements between the two sub-texts. For the above example, the statement combination with the highest similarity is selected from the 12 calculation results. Furthermore, the closest statements between any two sub-texts in each sub-class cluster are collected and organized to form a set containing all the closest statements. The intersection operation is performed on all the collected closest statements, that is, the statements that appear in all the closest statement combinations are found. By comparing the statements in different combinations multiple times, the common part can be gradually screened out, and thus the number of times each statement appears in the intersection result can be counted. The statement with the largest number of occurrences is determined as the general statement corresponding to the sub-class cluster. This general statement represents the core semantics shared by most sub-texts in the sub-class cluster.
[0071] Exemplarily, the position information where the general statement appears in each sub-text is obtained from the sub-class cluster, and then the position information is summarized to obtain the second position corresponding to the general statement. The second position can be the position information where the general statement appears the most times in the sub-text, or it can be a set of position information where the general statement appears in the sub-text.
[0072] In some embodiments, obtaining the target clustering result according to the second text includes: identifying keywords of the second text to obtain the text keywords corresponding to the second text; merging the text keywords to obtain all keywords, and calculating the similarity between any two keywords in the all keywords to obtain the relevant similarity value; classifying the all keywords according to the relevant similarity value to obtain a first phrase group and a second phrase group; performing cluster classification on the first phrase group by using the text keywords corresponding to the first text to obtain a first classification result, and obtaining the first central word corresponding to each first sub-classification result and the first central weight corresponding to the first central word according to the first classification result; performing cluster classification on the second phrase group by using the text keywords corresponding to the first text to obtain a second classification result, and obtaining the second central word corresponding to each second sub-classification result and the second central weight corresponding to the second central word according to the second classification result; determining the first associated word corresponding to the first central word in the second central word and the second associated word corresponding to the second central word in the first central word according to the first central word and the second central word in combination with the first central weight and the second central weight; matching the first classification result and the second classification result according to the first associated word and the second associated word to obtain the fusion similarity corresponding to multiple fusion phrase groups; calculating the text similarity corresponding to the second texts according to the fusion similarity and the fusion phrase groups; and performing text clustering on the second texts according to the text similarity to obtain the target clustering result.
[0073] Exemplarily, the second text is the first text after being processed by the target keywords. Then, keyword processing is performed on the second text according to the term frequency-inverse document frequency algorithm to identify the keywords therein, so as to obtain the text keywords corresponding to the second text, and all the identified text keywords are aggregated together, and duplicate words are removed to form all keywords.
[0074] Exemplarily, a similarity calculation method such as edit distance or cosine similarity is used to calculate the similarity between any two keywords in all keywords to obtain the relevant similarity value. Thus, the keywords with the relevant similarity value greater than the preset threshold are grouped into one category, and the keywords with the relevant similarity value less than or equal to the preset threshold are grouped into another category, so as to form a first phrase group and a second phrase group.
[0075] Exemplarily, based on the first phrase, combine with the text keywords corresponding to the second text, calculate the similarity between the text keywords and the first phrase, and then classify the second text into clusters according to the similarity result to obtain the first classification result. Similarly, based on the second phrase, combine with the text keywords corresponding to the second text, calculate the similarity between the text keywords and the second phrase, and then classify the second text into clusters according to the similarity result to obtain the second classification result.
[0076] Exemplarily, for each first sub-classification result in the first classification result, determine the first central word corresponding to the first sub-classification result by calculating the comprehensive importance of the keywords in this sub-classification, such as combining factors like word frequency and position in the sub-classification, and assign the corresponding first central weight according to its importance. Similarly, for each second sub-classification result in the second classification result, determine the second central word corresponding to the second sub-classification result by calculating the comprehensive importance of the keywords in this sub-classification, such as combining factors like word frequency and position in the sub-classification, and assign the corresponding second central weight according to its importance.
[0077] Exemplarily, compare the first central word and the second central word, combine the first central weight and the second central weight, and find the first associated word with a relatively high degree of association of the first central word in the second central word, as well as the second associated word corresponding to the second central word in the first central word.
[0078] Exemplarily, calculate the similarity based on the first associated word and the second associated word, so as to match the first sub-classification result in the first classification result and the second sub-classification result in the second classification result when the similarity is the largest, so that the first central word corresponding to the first sub-classification result and the second central word corresponding to the second sub-classification result form a fusion phrase. Then, after deleting the first associated word and the second associated word corresponding to the already formed fusion phrase, continue to obtain the maximum value corresponding to the remaining similarity calculation result, so as to form another corresponding fusion phrase, and then obtain multiple fusion phrases. Furthermore, for each fusion phrase, comprehensively consider the similarity between the included words and the corresponding weights, and then determine the corresponding fusion similarity.
[0079] Exemplarily, use a pre-trained language model to calculate the similarity between the fusion phrase and the semantic vector of the second text, and use this as the data association degree. Then, multiply the data association degree between each second text and each fusion phrase by the fusion similarity of this fusion phrase, and then sum up the calculation results of all fusion phrases. The obtained result is the text similarity corresponding to the second text under all fusion phrases. Thus, use the calculated text similarity as the input and use a clustering algorithm to cluster the second text to finally obtain the target clustering result.
[0080] Specifically, the process of determining associated words and fusion phrases helps to uncover the deep connections between texts. It can not only find surface-level lexical matches but also discover potential semantic links, providing more valuable information for further text clustering.
[0081] In some embodiments, calculating the corresponding text similarity between the second texts based on the fusion similarity and the fusion phrases includes: determining a first sub-text and a second sub-text from the second texts; obtaining a first frequency corresponding to each sub-keyword in the fusion phrase in the first sub-text and a second frequency corresponding to each sub-keyword in the second sub-text; obtaining a third sub-classification result corresponding to the fusion phrase from the first classification result and a fourth sub-classification result corresponding to the fusion phrase from the second classification result; obtaining a first quantity corresponding to the keywords in the third sub-classification result and a second quantity corresponding to the keywords in the fourth sub-classification result; fusing the fusion similarity corresponding to the fusion phrase based on the first frequency, the second frequency, the first quantity, and the second quantity to obtain the text similarity corresponding to the first sub-text and the second sub-text; wherein, the text similarity is obtained according to the following formula:
[0082]
[0083] Wherein, represents the text similarity corresponding to the i-th first sub-text and the j-th second sub-text, num2 represents the quantity corresponding to the fusion phrase, num1 represents the number of words corresponding to the sub-keyword in the q-th fusion phrase, represents the first frequency corresponding to the r-th sub-keyword in the q-th fusion phrase in the i-th first sub-text, represents the second frequency corresponding to the r-th sub-keyword in the q-th fusion phrase in the j-th second sub-text, represents the first quantity corresponding to the keywords in the third sub-classification result corresponding to the q-th fusion phrase, represents the second quantity corresponding to the keywords in the fourth sub-classification result corresponding to the q-th fusion phrase, represents the fusion similarity corresponding to the q-th fusion phrase.
[0084] Exemplarily, any two sub-texts are randomly selected from the second texts and defined as the first sub-text and the second sub-text respectively. Here, they can be selected sequentially or randomly, and the specific method is determined according to actual needs.
[0085] Exemplarily, for each sub-keyword in the fusion phrase, word frequency statistics are respectively performed in the first sub-text and the second sub-text. The first sub-text is scanned word by word to record the number of times each sub-keyword appears, obtaining the corresponding first frequency; similarly, the second sub-text is scanned to record the number of times each sub-keyword appears, obtaining the corresponding second frequency.
[0086] Exemplarily, according to the first classification result, the sub-classification result corresponding to the current fusion phrase is found therefrom and determined as the third sub-classification result. Similarly, in the second classification result, the sub-classification result corresponding to the fusion phrase is found and determined as the fourth sub-classification result.
[0087] Exemplarily, the keywords in the third sub-classification result are counted to obtain the first quantity; the keywords in the fourth sub-classification result are counted to obtain the second quantity.
[0088] Exemplarily, the first frequency, the second frequency, the first quantity, the second quantity obtained previously, and the fusion similarity corresponding to the fusion phrase are substituted into the following formula for calculation to obtain the text similarity:
[0089]
[0090] where represents the text similarity corresponding to the i-th first sub-text and the j-th second sub-text, num2 represents the quantity corresponding to the fusion phrase, num1 represents the number of words corresponding to the sub-keywords in the q-th fusion phrase, represents the first frequency corresponding to the r-th sub-keyword in the q-th fusion phrase in the i-th first sub-text, represents the second frequency corresponding to the r-th sub-keyword in the q-th fusion phrase in the j-th second sub-text represents the first quantity corresponding to the keywords in the third sub-classification result corresponding to the q-th fusion phrase, represents the second quantity corresponding to the keywords in the fourth sub-classification result corresponding to the q-th fusion phrase, represents the fusion similarity corresponding to the q-th fusion phrase.
[0091] Exemplarily, according to the above formula, text similarity calculation is performed. After the first sub-text and the second sub-text are deeply mined and processed, the extracted text information is contained in the fusion phrase and the sub-classification result. Thus, when calculating the text similarity between the first sub-text and the second sub-text, taking the proportion of the number of words in the fusion phrase to the total number of text keywords as the weight, the text similarity between the first sub-text and the second sub-text is obtained. The above method can reduce invalid calculations to a certain extent while considering the semantic information of words, and can better capture the similarity relationship between words.
[0092] Specifically, through the analysis of the fused phrases and sub-keywords, and by combining the information in different classification results, the semantic information of the text can be deeply mined. For example, the frequency of sub-keywords reflects their importance in the text, and the number of keywords in different classification results reflects the characteristics of the text under different classification systems. All these are helpful for better understanding the semantic connotation of the text, so as to provide a good basis for the subsequent text clustering with accurate text similarity. It can more accurately classify texts with similar semantics into one category, improve the quality of the clustering results, and make the clusters after clustering more representative and distinguishable.
[0093] In some embodiments, the obtaining of the general statement of the sub-cluster and the second position of the general statement by performing similar text analysis on each sub-cluster in the target clustering result includes: obtaining the corresponding third sub-text and fourth sub-text in the sub-cluster, and performing text segmentation on the third sub-text to obtain a first segmentation result and performing text segmentation on the fourth sub-text to obtain a second segmentation result; obtaining the corresponding third sub-statement in the first segmentation result and the corresponding fourth sub-statement in the second segmentation result; performing keyword recognition on the third sub-statement to obtain a first keyword and performing keyword recognition on the fourth sub-statement to obtain a second keyword; performing an intersection process on the first keyword and the second keyword to obtain the same keywords corresponding to the third sub-statement and the fourth sub-statement and the target number corresponding to the same keywords; performing a quantity statistics on the first keyword to obtain a third quantity, and performing a logarithm solution on the third quantity to obtain a first result; performing a quantity statistics on the second keyword to obtain a fourth quantity, and performing a logarithm solution on the fourth quantity to obtain a second result; performing a summation on the first result and the second result to obtain a target result, and performing a ratio calculation on the target number and the target result to obtain the corresponding statement similarity between the third sub-statement and the fourth sub-statement; determining the corresponding similar statement between the third sub-text and the fourth sub-text according to the statement similarity; performing data statistics on the similar statement according to the sub-cluster to obtain the general statement corresponding to the sub-cluster; and performing a position search on the general statement in the sub-cluster to obtain the second position corresponding to the general statement.
[0094] Exemplarily, two sub-texts are randomly selected from each sub-cluster of the target clustering result to obtain the corresponding third sub-text and fourth sub-text, and then the third sub-text is segmented into individual parts according to punctuation marks (period, exclamation mark, question mark, etc.), line breaks or specific delimiter identifiers to obtain a first segmentation result. Using the same method as the segmentation of the third sub-text, the fourth sub-text is segmented to obtain a second segmentation result.
[0095] Exemplarily, a third sub-statement is extracted from the first segmentation result, and a fourth sub-statement is extracted from the second segmentation result. The third sub-statement and the fourth sub-statement are usually short sentences with complete semantics. Thus, keyword recognition is used, such as word frequency statistics and graph-based sorting algorithms, to analyze the third sub-statement, find representative and important words therein, obtain the first keyword, and use the same keyword recognition method to process the fourth sub-statement to obtain the second keyword.
[0096] Exemplarily, the first keyword and the second keyword are compared to find the keywords they commonly contain. These common keywords are the same keywords corresponding to the third sub-statement and the fourth sub-statement, and the number of the same keywords is counted to obtain the target number. Furthermore, the number of the first keywords is counted to obtain the third number, and then the logarithm of the third number is taken to obtain the first result. The number of the second keywords is counted to obtain the fourth number, and then the logarithm of the fourth number is taken to obtain the second result.
[0097] Exemplarily, the first result and the second result are added to obtain the target result. Thus, the target number is divided by the target result to obtain the corresponding statement similarity between the third sub-statement and the fourth sub-statement.
[0098] Exemplarily, a similarity threshold is set. When the statement similarity between the third sub-statement and the fourth sub-statement is greater than the threshold, these two sub-statements are considered similar statements. Thus, all the third sub-statements in the third sub-text and all the fourth sub-statements in the fourth sub-text are calculated pairwise to obtain the similar statements between the third sub-text and the fourth sub-text.
[0099] Exemplarily, data statistics are performed on all the similar statements in the entire sub-cluster, such as counting the frequency of occurrence of each similar statement. The similar statements with a higher frequency of occurrence can be determined as the general statements corresponding to the sub-cluster.
[0100] Exemplarily, in all the texts of the sub-cluster, the positions where the general statements appear are searched one by one, and these position information are recorded to obtain the second position corresponding to the general statements.
[0101] Specifically, through meticulous text processing and similarity calculation, the similarities between the texts in the sub-cluster can be accurately found, and the general statements can be refined. These general statements reflect the fixed statements corresponding to the sub-cluster texts and are not related to the core events that the content of this official document needs to convey, thus providing good support for the subsequent generation of official document content with certain requirements or specifications.
[0102] Step S105: Perform keyword recognition on the general statement to obtain the second word and the second type of the second word.
[0103] Exemplarily, the named entity recognition model is a deep learning-based bidirectional long short-term memory network and conditional random field model, so as to use the named entity recognition model to perform sentence recognition on the general sentence to obtain the word type identifier corresponding to each word in the general sentence, and determine the second word corresponding to the general sentence and the second type corresponding to the second word according to the word type identifier.
[0104] Step S106: Extract a first sentence rule from the key sentence according to the first word and the first type.
[0105] Exemplarily, perform event classification on the key sentence through a machine learning classification algorithm to obtain the first event type corresponding to the key sentence, and then determine the first event association corresponding to the first type and the first event type according to the first word, so as to obtain a first rule after replacing the first word in the key sentence with the first type, and then determine the first sentence rule corresponding to the key sentence according to the first event association and the first rule between the first type and the first event type.
[0106] For example, extract features related to the event from the key sentence, such as the subject (person name, organization name, company name), action, time, location, etc. of the event. Thus, replace the first word in the key sentence with the corresponding first type, and mark the first type with square brackets or double quotes to distinguish it from other content. Then determine the relationship between the first type represented by the first word and the already determined event type. For example, if the first word is "school" and the event type is "education event", then it can be analyzed that the type "school" has a close association with the "education event", and the event relationship between the first type and the event type is determined to be the event subject. Thus, replace the first word in the key sentence with the first type. For example, if the first word is "school" and the first type is "organization name", the key sentence "The school prohibits random charging" becomes "
Organization name
[0107] Step S107: Extract a second sentence rule from the general sentence according to the second word and the second type.
[0108] Exemplarily, perform event classification on the general sentence through a machine learning classification algorithm to obtain the second event type corresponding to the general sentence, and then determine the second event association corresponding to the second type and the second event type according to the second word, so as to obtain a second rule after replacing the second word in the general sentence with the second type, and then determine the second sentence rule corresponding to the general sentence according to the first event association and the second rule between the second type and the second event type.
[0109] It should be noted that the general sentence is a common sentence corresponding to the official document content when it is released, and is often irrelevant to the event to be released, while the key sentence is the event information corresponding to the official document content when it is released.
[0110] Step S108: Obtain the target keyword of the event to be announced and the third type of the target keyword.
[0111] Exemplarily, obtain the event to be announced that the target user needs to publish, and then identify each word in the event to be announced according to the named entity recognition model to obtain the word type corresponding to each word in the event to be announced, so as to obtain the target keyword of the event to be announced and the third type of the target keyword according to the word type.
[0112] Step S109: Generate the target official document content of the event to be announced according to the target keyword and the third type, in combination with the first sentence rule and the second sentence rule, at the first position and the second position.
[0113] Exemplarily, calculate the similarity between the event to be announced and the event types corresponding to the first sentence rule and the second sentence rule, so as to screen out the first associated rule and the second associated rule corresponding to the event to be announced from the first sentence rule and the second sentence rule.
[0114] Exemplarily, determine the event relationship between the target keyword and the event to be announced, and then generate the corresponding target key sentence according to the event relationship and the third type by following the first associated rule for the target keyword, and generate the corresponding target general sentence according to the event relationship and the third type by following the second associated rule for the target keyword, so as to generate the target official document content of the event to be announced at the first position in the first text according to the first associated rule and at the second position in the first text according to the second associated rule.
[0115] In some embodiments, the step of generating the target official document content of the event to be announced according to the target keyword and the third type, in combination with the first sentence rule and the second sentence rule, at the first position and the second position includes: obtaining a first generated sentence according to the target keyword and the third type in combination with the first sentence rule; obtaining a second generated sentence according to the target keyword and the third type in combination with the second sentence rule; performing sentence merging on the first generated sentence and the second generated sentence according to the first position and the second position to obtain a plurality of initial official document contents; performing text quality evaluation on the initial official document contents according to the quality evaluation model to obtain the target text quality corresponding to the initial official document contents; and screening out the target official document content of the event to be announced from the initial official document contents according to the target text quality.
[0116] Exemplarily, calculate the similarity between the event to be announced and the event types corresponding to the first sentence rule and the second sentence rule, so as to screen out the first associated rule and the second associated rule corresponding to the event to be announced from the first sentence rule and the second sentence rule.
[0117] Exemplarily, determine the event relationship between the target keyword and the event to be announced, so as to generate the corresponding first generation statement for the target keyword according to the event relationship and the third type in combination with the first association rule, and generate the corresponding second generation statement for the target keyword according to the event relationship and the third type in combination with the second association rule.
[0118] Exemplarily, the first position and the second position indicate the arrangement order and combination method when the first generation statement and the second generation statement are merged. Therefore, according to the requirements of the first position and the second position, the first generation statement and the second generation statement are merged to obtain the initial merged content. Furthermore, with the help of a text generation model, text expansion is performed based on the initial merged content to obtain multiple initial official document contents. The text generation model can be a model based on a neural network or a model based on deep learning.
[0119] Exemplarily, construct a model that can comprehensively evaluate the text quality. This model can consider multiple factors, such as syntactic correctness, semantic coherence, logical rationality, information integrity, etc. Then, each initial official document content is input into the quality evaluation model, and it is scored according to the evaluation indicators and criteria of the model to obtain the target text quality corresponding to each initial official document content.
[0120] Exemplarily, set a threshold for the target text quality, so as to screen out the official document contents greater than this threshold from all the initial official document contents. Furthermore, the official document contents with a target text quality greater than this threshold are the target official document contents of the event to be announced.
[0121] Please refer to Figure 2 , Figure 2A document content generation device 200 based on artificial intelligence provided by an embodiment of the present application. The document content generation device 200 based on artificial intelligence includes a sentence recognition module 201, a first word recognition module 202, a replacement processing module 203, a clustering analysis module 204, a second word recognition module 205, a first rule extraction module 206, a second rule extraction module 207, a data acquisition module 208, and a document generation module 209. Among them, the sentence recognition module 201 is used to perform sentence recognition on the first text to obtain key sentences and the first positions of the key sentences; the first word recognition module 202 is used to perform keyword recognition on the key sentences to obtain first words and the first types of the first words; the replacement processing module 203 is used to perform word replacement on the first words in the first text according to the first type to obtain a second text; the clustering analysis module 204 is used to perform text clustering on the second text to obtain a target clustering result, and perform similar text analysis on each sub-cluster in the target clustering result to obtain general sentences of the sub-cluster and the second positions of the general sentences; the second word recognition module 205 is used to perform keyword recognition on the general sentences to obtain second words and the second types of the second words; the first rule extraction module 206 is used to perform rule extraction on the key sentences according to the first words and the first types to obtain a first sentence rule; the second rule extraction module 207 is used to perform rule extraction on the general sentences according to the second words and the second types to obtain a second sentence rule; the data acquisition module 208 is used to obtain target keywords of the event to be announced and the third types of the target keywords; the document generation module 209 is used to generate target document content of the event to be announced according to the target keywords and the third types in combination with the first sentence rule and the second sentence rule according to the first position and the second position.
[0122] In some embodiments, the document content generation device 200 based on artificial intelligence can be applied to a terminal device.
[0123] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the above-described document content generation device 200 based on artificial intelligence can refer to the corresponding process in the foregoing embodiment of the document content generation method based on artificial intelligence, and will not be described in detail here.
[0124] Please refer to Figure 3 , Figure 3 A schematic block diagram of the structure of a terminal device provided by an embodiment of the present invention.
[0125] Such as Figure 3As shown in the figure, the terminal device 300 includes a processor 301 and a memory 302. The processor 301 and the memory 302 are connected through a bus 303, which is, for example, an I2C (Inter-integrated Circuit) bus.
[0126] Specifically, the processor 301 is used to provide computing and control capabilities to support the operation of the entire terminal device. The processor 301 can be a central processing unit (CPU), or it can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0127] Specifically, the memory 302 can be a Flash chip, read-only memory (ROM), magnetic disk, optical disc, USB flash drive, or mobile hard disk, etc.
[0128] Those skilled in the art can understand that Figure 3 the structure shown in the figure is only a block diagram of some structures related to the solution of the embodiment of the present invention, and does not constitute a limitation on the terminal device to which the solution of the embodiment of the present invention is applied. A specific server may include more or fewer components than those shown in the figure, or combine certain components, or have a different component layout.
[0129] The processor is used to run a computer program stored in the memory and, when executing the computer program, implement any one of the artificial intelligence-based official document content generation methods provided by the embodiments of the present invention.
[0130] In one embodiment, the processor is used to run a computer program stored in the memory and, when executing the computer program, implement the following steps:
[0131] Perform sentence recognition on the first text to obtain a key sentence and the first position of the key sentence;
[0132] Perform keyword recognition on the key sentence to obtain a first word and the first type of the first word;
[0133] Perform word substitution on the first word in the first text according to the first type to obtain a second text;
[0134] Perform text clustering on the second text to obtain a target clustering result, and perform similar text analysis on each sub-cluster in the target clustering result to obtain a general statement of the sub-cluster and a second position of the general statement;
[0135] Perform keyword recognition on the general statement to obtain a second word and a second type of the second word;
[0136] Extract a first statement rule from the key sentence according to the first word and the first type;
[0137] Extract a second statement rule from the general statement according to the second word and the second type;
[0138] Obtain a target keyword of the event to be announced and a third type of the target keyword;
[0139] Generate a target official document content of the event to be announced according to the target keyword and the third type in combination with the first statement rule and the second statement rule according to the first position and the second position.
[0140] It should be noted that those skilled in the art can clearly understand that, for the convenience and conciseness of description, the specific working process of the above-described terminal device can refer to the corresponding process in the foregoing embodiment of the official document content generation method based on artificial intelligence, and will not be described in detail here.
[0141] The embodiment of the present invention further provides a storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of any one of the official document content generation methods provided in the specification of the embodiment of the present invention.
[0142] Among them, the storage medium may be an internal storage unit of the terminal device described in the foregoing embodiment, such as the hard disk or memory of the terminal device. The storage medium may also be an external storage device of the terminal device, such as a plug-in hard disk equipped on the terminal device, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.
[0143] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations. In a hardware embodiment, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, one physical component can have multiple functions, or one function or step can be executed by several physical components in cooperation. Some or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory, or other memory technologies, CD-ROM, digital versatile disks (DVDs), or other optical disk storage, magnetic cartridges, tapes, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0144] It should be understood that the term "and / or" used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. It should be noted that in this text, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or system comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements that are inherent to such process, method, article, or system. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article, or system comprising that element.
[0145] The serial numbers of the embodiments of the present invention above are only for description and do not represent the advantages or disadvantages of the embodiments. As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for generating official document content based on artificial intelligence, characterized in that, The method includes: Performing sentence recognition on the first text to obtain a key sentence and a first position of the key sentence; Performing keyword recognition on the key sentence to obtain a first word and a first type of the first word; Performing word replacement on the first word in the first text according to the first type to obtain a second text; Performing text clustering according to the second text to obtain a target clustering result, and performing similar text analysis on each sub-cluster in the target clustering result to obtain a general sentence of the sub-cluster and a second position of the general sentence; Performing keyword recognition on the general sentence to obtain a second word and a second type of the second word; Performing rule extraction on the key sentence according to the first word and the first type to obtain a first sentence rule; Performing rule extraction on the general sentence according to the second word and the second type to obtain a second sentence rule; Obtaining a target keyword of the event to be announced and a third type of the target keyword; Generating target official document content of the event to be announced according to the target keyword and the third type in combination with the first sentence rule and the second sentence rule according to the first position and the second position; Among them, the performing sentence recognition on the first text to obtain a key sentence and a first position of the key sentence includes: Performing keyword extraction on the first text to obtain a third word corresponding to the first text and a word weight corresponding to the third word; Performing sentence segmentation on the first text to obtain initial sentences, and determining a sentence weight corresponding to the initial sentences according to the third word and the word weight; Determining any one sentence in the initial sentences as a first sentence, and obtaining a second sentence corresponding to the first sentence in an adjacent state from the initial sentences; Obtaining a first weight corresponding to the first sentence and a second weight corresponding to the second sentence from the sentence weights; Performing weight adjustment on the first weight according to the second weight to obtain a target weight corresponding to the first sentence; Performing sentence recognition on the initial sentences according to the target weight to obtain the key sentence, and obtaining the first position corresponding to the key sentence from the first text according to the key sentence.
2. The method according to claim 1, characterized in that, The performing weight adjustment on the first weight according to the second weight to obtain a target weight corresponding to the first sentence includes: Obtaining a first sequence of sentences corresponding to the position on the left side of the second sentence from the initial sentences, and obtaining a second sequence of sentences corresponding to the position on the right side of the second sentence from the initial sentences; Obtaining an associated keyword corresponding to each first sub-sentence in the first sequence of sentences from the third word, and obtaining a relevant weight corresponding to the associated keyword from the word weights; Obtaining a third weight corresponding to each first sub-sentence in the first sequence of sentences from the sentence weights and obtaining a fourth weight corresponding to each second sub-sentence in the second sequence of sentences from the sentence weights; Performing weight adjustment on the first weight according to the second weight by combining the relevant weight, the third weight and the fourth weight to obtain the target weight corresponding to the first sentence; Among them, the target weight is obtained according to the following formula: Among them, represents the target weight corresponding to the i-th first statement, represents the adjustment parameter, represents the first weight corresponding to the i-th first statement, n represents the number of sentences corresponding to the first sequence of statements, and y represents the number of associated keywords corresponding to the h-th first sub-statement, represents the relevant weight corresponding to the k-th associated keyword corresponding to the h-th first sub-statement, represents the third weight corresponding to the h-th first sub-statement, and g represents the number of sentences corresponding to the second sequence of statements, represents the fourth weight corresponding to the t-th second sub-statement, represents the second weight of the second statement corresponding to the adjacent state of the i-th first statement.
3. The method according to claim 1, wherein The text clustering according to the second text to obtain a target clustering result includes: Performing keyword recognition on the second text to obtain text keywords corresponding to the second text; Performing word combination on the text keywords to obtain all keywords, and calculating the similarity between any two keywords in the all keywords to obtain relevant similarity values; Classifying the all keywords according to the relevant similarity values to obtain a first phrase group and a second phrase group; Performing cluster classification on the first phrase group by using the text keywords corresponding to the second text to obtain a first classification result, and obtaining a first central word corresponding to each first sub-classification result and a first central weight corresponding to the first central word according to the first classification result; Performing cluster classification on the second phrase group by using the text keywords corresponding to the second text to obtain a second classification result, and obtaining a second central word corresponding to each second sub-classification result and a second central weight corresponding to the second central word according to the second classification result; Combining the first central word and the second central word with the first central weight and the second central weight to determine a first associated word corresponding to the first central word in the second central word and a second associated word corresponding to the second central word in the first central word; Matching the first classification result and the second classification result according to the first associated word and the second associated word to obtain fusion similarity values corresponding to multiple fusion phrase groups; Calculating the text similarity corresponding to the second texts according to the fusion similarity values and the fusion phrase groups; Performing text clustering on the second texts according to the text similarity to obtain the target clustering result.
4. The method according to claim 3, wherein The calculating the text similarity corresponding to the second texts according to the fusion similarity values and the fusion phrase groups includes: Determining a first sub-text and a second sub-text from the second texts; Obtaining a first frequency corresponding to each sub-keyword in the fusion phrase group in the first sub-text and a second frequency corresponding to each sub-keyword in the second sub-text; Obtaining a third sub-classification result corresponding to the fusion phrase group from the first classification result and obtaining a fourth sub-classification result corresponding to the fusion phrase group from the second classification result; Obtaining a first quantity corresponding to keywords in the third sub-classification result and obtaining a second quantity corresponding to keywords in the fourth sub-classification result; Fusing the fusion similarity values corresponding to the fusion phrase group according to the first frequency, the second frequency, the first quantity and the second quantity to obtain the text similarity corresponding to the first sub-text and the second sub-text; Among them, the text similarity is obtained according to the following formula: Among them, represents the text similarity corresponding to the i-th first sub-text and the j-th second sub-text, num2 represents the quantity corresponding to the fusion phrase, and num1 represents the number of words corresponding to the sub-keyword in the q-th fusion phrase. represents the first frequency corresponding to the r-th sub-keyword in the q-th fusion phrase in the i-th first sub-text. represents the second frequency corresponding to the r-th sub-keyword in the q-th fusion phrase in the j-th second sub-text. represents the first quantity corresponding to the keyword in the third sub-classification result corresponding to the q-th fusion phrase. represents the second quantity corresponding to the keyword in the fourth sub-classification result corresponding to the q-th fusion phrase. represents the fusion similarity corresponding to the q-th fusion phrase.
5. The method according to claim 1, wherein The performing similar text analysis on each sub-cluster in the target clustering result to obtain a general statement of the sub-cluster and a second position of the general statement includes: Obtain the corresponding third sub-text and fourth sub-text in the sub-class cluster, perform text segmentation on the third sub-text to obtain a first segmentation result, and perform text segmentation on the fourth sub-text to obtain a second segmentation result; Obtain the corresponding third sub-statement in the first segmentation result and the corresponding fourth sub-statement in the second segmentation result; Perform keyword recognition on the third sub-statement to obtain a first keyword and perform keyword recognition on the fourth sub-statement to obtain a second keyword; Perform an intersection process on the first keyword and the second keyword to obtain the same keyword corresponding to the third sub-statement and the fourth sub-statement and the target quantity corresponding to the same keyword; Perform a quantity statistics on the first keyword to obtain a third quantity, and perform a logarithm solution on the third quantity to obtain a first result; Perform a quantity statistics on the second keyword to obtain a fourth quantity, and perform a logarithm solution on the fourth quantity to obtain a second result; Sum the first result and the second result to obtain a target result, and calculate the ratio of the target quantity and the target result to obtain the statement similarity corresponding to the third sub-statement and the fourth sub-statement; Determine the corresponding similar statements between the third sub-text and the fourth sub-text according to the statement similarity; Perform data statistics on the similar statements according to the sub-class cluster to obtain the general statement corresponding to the sub-class cluster; Perform a position search in the sub-class cluster according to the general statement to obtain the second position corresponding to the general statement.
6. The method according to claim 1, characterized in that, The generating the target official document content of the to-be-announced event according to the target keyword and the third type in combination with the first statement rule and the second statement rule according to the first position and the second position includes: Obtain a first generated statement according to the target keyword and the third type in combination with the first statement rule; Obtain a second generated statement according to the target keyword and the third type in combination with the second statement rule; Perform statement merging on the first generated statement and the second generated statement according to the first position and the second position to obtain multiple initial official document contents; Perform text quality evaluation on the initial official document content according to the quality evaluation model to obtain the target text quality corresponding to the initial official document content; Screen the target official document content of the to-be-announced event from the initial official document content according to the target text quality.
7. An official document content generation device based on artificial intelligence, characterized in that, Including: A statement recognition module, configured to perform statement recognition on a first text to obtain a key sentence and a first position of the key sentence. Wherein, performing statement recognition on the first text to obtain the key sentence and the first position of the key sentence includes: extracting keywords from the first text to obtain a third word corresponding to the first text and a word weight corresponding to the third word; segmenting the first text into initial statements, and determining a statement weight corresponding to the initial statement according to the third word and the word weight; determining any one sentence in the initial statement as a first statement, and obtaining a second statement corresponding to the first statement in an adjacent state from the initial statement; obtaining a first weight corresponding to the first statement and a second weight corresponding to the second statement from the statement weights; adjusting the first weight according to the second weight to obtain a target weight corresponding to the first statement; performing statement recognition on the initial statement according to the target weight to obtain the key sentence, and obtaining the first position corresponding to the key sentence from the first text according to the key sentence; A first word recognition module, configured to perform keyword recognition on the key sentence to obtain a first word and a first type of the first word; A replacement processing module, configured to perform word replacement on the first word in the first text according to the first type to obtain a second text; A clustering analysis module, configured to perform text clustering on the second text to obtain a target clustering result, and perform similar text analysis on each sub-cluster in the target clustering result to obtain a general statement of the sub-cluster and a second position of the general statement; A second word recognition module, configured to perform keyword recognition on the general statement to obtain a second word and a second type of the second word; A first rule extraction module, configured to perform rule extraction on the key sentence according to the first word and the first type to obtain a first statement rule; A second rule extraction module, configured to perform rule extraction on the general statement according to the second word and the second type to obtain a second statement rule; A data acquisition module, configured to obtain a target keyword of an event to be announced and a third type of the target keyword; A document generation module, configured to generate target document content of the event to be announced according to the target keyword and the third type, in combination with the first statement rule and the second statement rule, according to the first position and the second position; 8. A terminal device, characterized in that, The terminal device includes a processor and a memory; The memory is used to store a computer program; The processor is configured to execute the computer program and, when executing the computer program, implement the artificial intelligence-based document content generation method according to any one of claims 1 to 6.
9. A computer storage medium for computer storage, characterized in that, The computer storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the artificial intelligence-based document content generation method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Text classification method and device, electronic equipment and readable storage medium
CN112597312A
Text intention classification method and device, equipment and storage medium
CN114860942A