Artificial intelligence-based official document content generation method and device, terminal equipment and storage medium

Through the artificial intelligence-based official document content generation method, the problem of low quality of official document generation in the existing technology is solved, standardized and rigorous official document content generation is achieved, and work efficiency is improved.

CN120012729AActive Publication Date: 2025-05-16ZHUHAI MEDIA SUNAC TECH CO LTD

Patent Information

Application Number
CN202510502214.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-16
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

When generating official document content, it is difficult for the prior art to ensure format specifications and language rigor, resulting in low quality of the generated official document and difficult to meet actual work needs.

Method used

Through artificial intelligence-based methods, including statement recognition, keyword recognition, word replacement, text clustering and rule extraction, the target official document content is generated to ensure the standardization and rigor of its format and language.

Benefits of technology

It has achieved the generation of relatively standardized and consistent official document content, improved the quality of official document generation, and met the needs of assisting actual work.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012729A_ABST
    Figure CN120012729A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an official document content generation method and device based on artificial intelligence, terminal equipment and a storage medium. The method comprises the steps of performing statement recognition on a first text to obtain a key sentence and a first position; identifying the key sentence to obtain a first word and a first type; performing word replacement on a first word in the first text according to the first type to obtain a second text; performing clustering according to the second text to obtain a clustering result, and performing similar text analysis on the clustering result to obtain a general statement and a second position thereof; identifying the general statement to obtain a second word and a second type; performing rule extraction on the key sentence according to the first word and the first type to obtain a first rule; performing rule extraction on the general statement according to the second word and the second type to obtain a second rule; obtaining a target keyword and a third type of the to-be-published event; and generating target official document content according to the first position and the second position in combination with the first rule and the second rule according to the target keyword and the third type.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an artificial intelligence-based document content generation method, device, terminal equipment and storage medium. Background Art

[0002] In actual application scenarios, taking government agencies and enterprises and institutions as examples, official documents are extremely important tools for uploading and recording matters. Official documents carry key functions such as information transmission and decision-making deployment, and have strict format specifications and rigorous language expression requirements. However, if the content of official documents is generated with the help of a text generation model, the format of the generated official documents may deviate from the specifications due to the randomness of the text generation model, and it may be difficult to meet the standards in aspects such as page settings and typesetting layout; it is also difficult to match the rigor of the official documents in terms of language expression, and there may be problems such as inappropriate wording and casual expressions. Therefore, when the existing technology generates text for official document content, the generated content is inconsistent with the actual needs, which leads to the low quality of the generated official documents, which is difficult to meet the needs of assisting actual work. Summary of the invention

[0003] The main purpose of the embodiments of the present invention is to provide a method, device, terminal device and storage medium for generating official document content based on artificial intelligence, aiming to solve the problem in the related technology that the quality of the official document generated when the text of the official document content is generated is low and it is difficult to meet the needs of assisting actual work.

[0004] In a first aspect, an embodiment of the present invention provides a method for generating official document content based on artificial intelligence, comprising:

[0005] Performing sentence recognition on the first text to obtain a key sentence and a first position of the key sentence;

[0006] Performing keyword recognition on the key sentence to obtain a first word and a first type of the first word;

[0007] Performing word replacement on the first word in the first text according to the first type to obtain a second text;

[0008] Performing text clustering according to the second text to obtain a target clustering result, and performing similar text analysis on each sub-cluster in the target clustering result to obtain a common sentence of the sub-cluster and a second position of the common sentence;

[0009] Performing keyword recognition on the general sentence to obtain a second word and a second type of the second word;

[0010] Extracting rules from the key sentence according to the first word and the first type to obtain a first sentence rule;

[0011] Extracting rules from the general sentence according to the second word and the second type to obtain a second sentence rule;

[0012] Obtaining a target keyword of an event to be announced and a third type of the target keyword;

[0013] The target official document content of the event to be announced is generated according to the target keyword and the third type in combination with the first sentence rule and the second sentence rule according to the first position and the second position.

[0014] In a second aspect, an embodiment of the present invention provides an official document content generation device based on artificial intelligence, comprising:

[0015] A sentence recognition module, used for performing sentence recognition on the first text to obtain a key sentence and a first position of the key sentence;

[0016] A first word recognition module, configured to perform keyword recognition on the key sentence to obtain a first word and a first type of the first word;

[0017] a replacement processing module, configured to replace the first word in the first text according to the first type to obtain a second text;

[0018] A cluster analysis module, configured to perform text clustering according to the second text to obtain a target clustering result, and perform similar text analysis on each sub-cluster in the target clustering result to obtain a common sentence of the sub-cluster and a second position of the common sentence;

[0019] A second word recognition module, used for performing keyword recognition on the general sentence to obtain a second word and a second type of the second word;

[0020] A first rule extraction module, configured to extract rules from the key sentence according to the first word and the first type to obtain a first sentence rule;

[0021] A second rule extraction module, configured to extract rules from the general sentence according to the second word and the second type to obtain a second sentence rule;

[0022] A data acquisition module, used to obtain a target keyword of an event to be announced and a third type of the target keyword;

[0023] The official document generation module is used to generate the target official document content of the event to be announced according to the first position and the second position in combination with the first sentence rule and the second sentence rule according to the target keyword and the third type.

[0024] In a third aspect, an embodiment of the present invention further provides a terminal device, comprising a processor, a memory, a computer program stored in the memory and executable by the processor, and a data bus for realizing connection and communication between the processor and the memory, wherein when the computer program is executed by the processor, the steps of any one of the methods for generating official document content based on artificial intelligence provided in the specification of the present invention are realized.

[0025] In a fourth aspect, an embodiment of the present invention further provides a storage medium for computer-readable storage, characterized in that the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of any one of the methods for generating official document content based on artificial intelligence provided in the specification of the present invention.

[0026] The embodiment of the present invention provides a method, apparatus, terminal device and storage medium for generating official document content based on artificial intelligence. The method comprises: performing sentence recognition on a first text to obtain a key sentence and a first position of the key sentence, so as to quickly locate the core content corresponding to the first text and provide good support for subsequent rule extraction; performing keyword recognition on the key sentence to obtain a first word and a first type of the first word; performing word replacement on the first word in the first text according to the first type to obtain a second text, so as to achieve different corresponding expressions between the same type under different words, thereby reducing the interference of different words and providing good support for subsequent text clustering, and then performing text clustering according to the second text to obtain a target clustering result, and performing similar text analysis on each subclass cluster in the target clustering result to obtain a common sentence of the subclass cluster and a common sentence of the common sentence. The second position, and then perform keyword recognition on the general sentence to obtain the second word and the second type of the second word, thereby performing rule extraction on the key sentence according to the first word and the first type to obtain the first sentence rule; and performing rule extraction on the general sentence according to the second word and the second type to obtain the second sentence rule; obtain the target keyword of the event to be announced and the third type of the target keyword; generate the target official document content of the event to be announced according to the first position and the second position according to the target keyword and the third type combined with the first sentence rule and the second sentence rule, thereby generating more standardized and consistent official document content, further improving the quality of generated official documents, providing support for further improving the work efficiency of relevant personnel, and also solving the problem in the related technology that the quality of the official documents generated when performing text generation on the official document content is low and it is difficult to meet the needs of assisting actual work. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0028] Figure 1 A flowchart of a method for generating official document content based on artificial intelligence provided by an embodiment of the present invention;

[0029] Figure 2 A schematic diagram of the module structure of another device for generating official document content based on artificial intelligence provided by an embodiment of the present invention;

[0030] Figure 3 A schematic block diagram of the structure of a terminal device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0031] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0032] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may also be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions.

[0033] It should be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.

[0034] The embodiment of the present invention provides a method, device, terminal device and storage medium for generating official document content based on artificial intelligence. The method for generating official document content based on artificial intelligence can be applied to a terminal device, which can be an electronic device such as a tablet computer, a laptop computer, a desktop computer, a personal digital assistant and a wearable device. The terminal device can be a server or a server cluster.

[0035] Some embodiments of the present invention are described in detail below in conjunction with the accompanying drawings. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.

[0036] Please refer to Figure 1 , Figure 1 A flowchart of a method for generating official document content based on artificial intelligence is provided in accordance with an embodiment of the present invention.

[0037] like Figure 1 As shown, the document content generation method based on artificial intelligence includes steps S101 to S109.

[0038] Step S101: Perform sentence recognition on a first text to obtain a key sentence and a first position of the key sentence.

[0039] Exemplarily, the first text is a historical official document text that has been published or announced obtained from a database, and then a target title is obtained from the historical official document text, and the historical official document text is segmented to obtain multiple segmented sentences, and then the degree of association or similarity between each segmented sentence and the target title is judged according to the cosine similarity, so that the segmented sentence corresponding to the maximum degree of association or similarity is determined as the corresponding key sentence.

[0040] Exemplarily, after the historical official document text is segmented, the sentence number corresponding to each segmented sentence is obtained, and then when the segmented sentence corresponding to the maximum correlation or similarity is determined as the key sentence, the sentence number corresponding to the segmented sentence is determined as the first position corresponding to the key sentence.

[0041] In some embodiments, the sentence recognition of the first text to obtain the key sentence and the first position of the key sentence includes: performing keyword extraction on the first text to obtain the third word corresponding to the first text and the word weight corresponding to the third word; performing sentence segmentation on the first text to obtain an initial sentence, and determining the sentence weight corresponding to the initial sentence according to the third word and the word weight; determining any sentence in the initial sentence as the first sentence, and obtaining the second sentence corresponding to the first sentence in an adjacent state from the initial sentence; obtaining the first weight corresponding to the first sentence and the second weight corresponding to the second sentence from the sentence weight; weight-adjusting the first weight according to the second weight to obtain the target weight corresponding to the first sentence; performing sentence recognition on the initial sentence according to the target weight to obtain the key sentence, and obtaining the first position corresponding to the key sentence from the first text according to the key sentence.

[0042] Exemplarily, a word frequency-inverse document frequency algorithm is used to perform keyword extraction on the first text to obtain a third word corresponding to the first text, and the word frequency corresponding to the third word is obtained from the first text, and then the word weight corresponding to the third word is determined based on the word frequency, wherein the word weight can be further adjusted according to the position where the third word appears.

[0043] Exemplarily, the first text is divided into a plurality of initial sentences according to sentence separators such as a period, a question mark, an exclamation mark, etc. Thus, for each initial sentence, the word weights of the third words contained in the initial sentence are accumulated to obtain the sentence weight of the initial sentence. For example, an initial sentence contains three keywords with word weights of 0.3, 0.2, and 0.1, then the sentence weight of the initial sentence is 0.3+0.2+0.1=0.6.

[0044] Exemplarily, a sentence is selected from a plurality of initial sentences in order as the first sentence. For example, the first initial sentence is selected as the first sentence. If the first sentence is the first sentence in the first text, then its adjacent second sentence is the sentence that follows it immediately; if the first sentence is not the first sentence, then its adjacent second sentence can be the sentence before it or the sentence after it, which is determined according to actual needs.

[0045] Exemplarily, after determining the first sentence and the second sentence, the first weight corresponding to the first sentence and the second weight corresponding to the second sentence are obtained from the sentence weights, so as to obtain the proportional coefficient by calculating the proportional relationship between the second weight and the first weight. When the proportional coefficient is greater than or equal to the preset value, it indicates that the gap between the first weight and the second weight is large, and the first weight can be directly determined as the target weight without processing the first weight according to the second weight, that is, the importance of the first sentence and the second sentence has been clearly distinguished. If additional weight adjustments are made, the original reasonable weight relationship may be destroyed, resulting in deviations in key sentence recognition. When the proportional coefficient is less than the preset value, it indicates that the gap between the first weight and the second weight is small, and the gap value between the first weight and the second weight is first calculated to determine an adjustment coefficient. The smaller data in the first weight and the second weight are adjusted downward, and the larger data in the first weight and the second weight are adjusted upward, that is, when the gap between the first weight and the second weight is small, the smaller weight is adjusted downward, and the larger weight is adjusted upward. In this way, the gap between the two can be further widened, so that the weight of the important sentence is more prominent, so that in the subsequent key sentence recognition process, it is easier to screen out the truly important sentences, so as to obtain the target weight corresponding to the adjusted first sentence.

[0046] Exemplarily, all initial sentences are respectively weighted as first sentences to obtain a target weight corresponding to each initial sentence, and a screening threshold is determined, and then the initial sentences with target weights greater than the threshold are determined as key sentences.

[0047] Exemplarily, the sentence sequence number corresponding to each initial sentence after sentence segmentation of the first text is obtained, and then the sentence sequence number corresponding to the key sentence is obtained from the sentence sequence number, so as to determine the sentence sequence number as the first position corresponding to the key sentence.

[0048] In some embodiments, the weight adjustment of the first weight according to the second weight to obtain the target weight corresponding to the first sentence includes: obtaining a first sequence sentence corresponding to the left position of the second sentence from the initial sentence, and obtaining a second sequence sentence corresponding to the right position of the second sentence from the initial sentence; obtaining an associated keyword corresponding to each first sub-sentence in the first sequence sentence from the third word, and obtaining a relevant weight corresponding to the associated keyword from the word weight; obtaining a third weight corresponding to each first sub-sentence in the first sequence sentence from the sentence weight and obtaining a fourth weight corresponding to each second sub-sentence in the second sequence sentence from the sentence weight; weight adjustment of the first weight using the second weight according to the relevant weight combined with the third weight and the fourth weight to obtain the target weight corresponding to the first sentence; wherein the target weight is obtained according to the following formula:

[0049]

[0050] in, represents the target weight corresponding to the i-th first sentence, represents the adjustment parameter, represents the first weight corresponding to the i-th first sentence, n represents the number of sentences corresponding to the first sequence of sentences, y represents the number of associated keywords corresponding to the h-th first sub-sentence, represents the relevant weight corresponding to the kth associated keyword corresponding to the hth first sub-sentence, represents the third weight corresponding to the hth first sub-sentence, g represents the number of sentences corresponding to the second sequence of sentences, represents the fourth weight corresponding to the t-th second sub-statement, Represents the second weight of the second sentence corresponding to the i-th adjacent state of the first sentence.

[0051] Exemplarily, after performing data segmentation on the first text to obtain multiple initial sentences, the multiple initial sentences are stored in the order in the first text to obtain a data storage result, so that after determining the second sentence corresponding to the first sentence, the position of the second sentence is found from the data storage result, and then all sentences on the left side of the second sentence are extracted to form a first sequence of sentences; all sentences on the right side of the second sentence are extracted to form a second sequence of sentences.

[0052] Exemplarily, the associated keywords associated with each first sub-sentence in the first sequence of sentences are found through word matching according to the third word, and then the relevant weights corresponding to these associated keywords are found according to the corresponding relationship between the third word and the word weight.

[0053] Exemplarily, according to the correspondence between the initial statement and the statement weight, the third weight corresponding to each first sub-statement in the first sequence of statements and the fourth weight corresponding to each second sub-statement in the second sequence of statements are respectively obtained from the statement weight.

[0054] Exemplarily, the target weight corresponding to the first sentence is obtained by weight-adjusting the first weight using the second weight according to the relevant weight combined with the third weight and the fourth weight using the following formula;

[0055]

[0056] in, represents the target weight corresponding to the i-th first sentence, represents the adjustment parameter, represents the first weight corresponding to the i-th first sentence, n represents the number of sentences corresponding to the first sequence sentence, y represents the number of associated keywords corresponding to the h-th first sub-sentence, represents the relevant weight of the kth associated keyword corresponding to the hth first sub-sentence, represents the third weight corresponding to the hth first sub-sentence, g represents the number of sentences corresponding to the second sequence sentence, represents the fourth weight corresponding to the t-th second sub-statement, Represents the second weight of the second sentence corresponding to the adjacent state of the i-th first sentence.

[0057] Exemplarily, the adjustment parameter is a pre-set value used to control the degree of adjustment; thereby, the complete information corresponding to the first sentence and the first text is fused through the weight information of the left sentence sequence and the right sentence sequence of the second sentence and the relevant weights corresponding to the associated keywords with the help of the second weight corresponding to the second sentence, so as to more accurately obtain the target weight corresponding to the first sentence, thereby providing better support for the subsequent acquisition of key sentences.

[0058] Specifically, when calculating the target weight corresponding to the first sentence, the first sentence is integrated with the complete information corresponding to the first text with the help of the second sentence, thereby making the weight adjustment more comprehensive and flexible, thereby further improving the reliability and accuracy of the weight calculation.

[0059] Step S102: performing keyword recognition on the key sentence to obtain a first word and a first type of the first word.

[0060] Exemplarily, the naming recognition model is used to recognize the key sentence to obtain the word type identifier corresponding to each word in the key sentence, and the first word corresponding to the key sentence and the first type corresponding to the first word are determined according to the word type identifier.

[0061] For example, the naming recognition model is a bidirectional long short-term memory network and conditional random field model based on deep learning. Therefore, the pre-trained model is fine-tuned using a labeled data set, wherein the data set contains a large number of sentences, and each word in the sentence is labeled with the corresponding word type, such as a person's name, a place name, an organization name, a date, and others. Therefore, the parameters of the model are continuously adjusted during the training process to improve the model's recognition accuracy of the word type, and then the key sentence is input into the trained naming recognition model. The model will analyze each word and output the corresponding word type identifier. Therefore, when the word type is identified as a person's name, a place name, an organization name, or a date, the corresponding word is determined as the first word, and the word type identifier corresponding to the first word is determined as the first type.

[0062] Step S103: Replace the first word in the first text according to the first type to obtain a second text.

[0063] Exemplarily, the first type includes person names, place names, organization names, time, etc., and then a common target replacement word is defined for each first type. For example, for the person name type, it can be uniformly replaced with "someone"; for the place name type, it can be replaced with "somewhere", so that the defined different first types and their corresponding target replacement words are organized into a replacement word library to facilitate subsequent search and use.

[0064] Exemplarily, the aforementioned naming recognition model is used to process the first text, identify the first word and the first type corresponding to the first word, and then search for the corresponding target replacement word in the replacement word library according to the first type, and then replace the corresponding first word in the first text according to the target replacement word to obtain the replaced second text. After the first word corresponding to the first type is replaced with a general expression, the second text can be applied to a wider range of scenarios, providing good support for subsequent text clustering.

[0065] For example, for the first text "Zhang San went to Beijing yesterday", through named entity recognition, we can determine that "Zhang San" is a person's name and "Beijing" is a place name, and then find the corresponding target replacement word from the replacement vocabulary. The target replacement word corresponding to the person's name is "someone", and the target replacement word corresponding to the place name is "somewhere", and then we can get the second text after replacement as "someone went to somewhere yesterday".

[0066] Step S104: performing text clustering according to the second text to obtain a target clustering result, and performing similar text analysis on each sub-cluster in the target clustering result to obtain a common sentence of the sub-cluster and a second position of the common sentence.

[0067] Exemplarily, using word embedding such as Word2Vec, GloVe, etc., the corresponding words in the second text are mapped to a low-dimensional vector space to obtain the word embedding vector corresponding to each word in the second text, and then the second text is represented by calculating the average or weighted average of the word embedding vectors of the words in the second text. After each second text is converted into a feature vector, these vectors are combined into a feature matrix, each row of the matrix represents a text, and each column represents a feature, so that a clustering algorithm such as hierarchical clustering is used to perform text clustering according to the feature matrix to obtain the target clustering result.

[0068] For example, since different types of words have been unified by using common words in the second text, the key words have been unified when the target clustering result is obtained by clustering the second text. Based on this, the texts contained in each sub-cluster in the target clustering result often have the same or similar meanings. In view of this, when similar text analysis is performed on each sub-cluster in the target clustering result, the common sentences corresponding to each sub-cluster can be summarized.

[0069] Exemplarily, the similarity of sentences between any two sub-texts in the corresponding sub-texts in each sub-class cluster in the target clustering result is calculated to obtain the most similar sentence between the two sub-texts, and then all the most similar sentences between any two sub-texts in the sub-class cluster are obtained and intersection processing is performed on all the most similar sentences, and then the sentences with the largest number in the intersection processing results are determined as the common sentences corresponding to the sub-class cluster.

[0070] For example, for any two subtexts in each subclass cluster, the similarity of each sentence in one subtext is calculated with all sentences in the other subtext. For example, if subtext A has 3 sentences and subtext B has 4 sentences, then 3×4 = 12 similarity calculations are required, so that in each group of sentence similarity calculation results, the sentence pairs with the highest similarity are found. These sentence pairs are the most similar sentences between the two subtexts. For the above example, the sentence combination with the highest similarity is selected from the 12 calculation results, and then the most similar sentences between any two subtexts in each subclass cluster are collected and sorted to form a set containing all the most similar sentences. Intersection processing is performed on all the collected most similar sentences, that is, to find the sentences that appear in all the most similar sentence combinations. By comparing sentences of different combinations multiple times, the common parts can be gradually screened out, so as to count the number of occurrences of each sentence in the intersection processing results. The sentence with the largest number of occurrences is determined as the common sentence corresponding to the subclass cluster. This common sentence represents the core semantics shared by most subtexts in the subclass cluster.

[0071] Exemplarily, the position information of the common sentence in each sub-text is obtained from the sub-class cluster, and then the position information is summarized to obtain the second position corresponding to the common sentence. The second position can be the position information of the common sentence appearing the most times in the sub-text, or it can be the collection of each position information of the common sentence appearing in the sub-text.

[0072] In some embodiments, the text clustering according to the second text to obtain the target clustering result includes: performing keyword recognition on the second text to obtain text keywords corresponding to the second text; performing word merging on the text keywords to obtain all keywords, and performing similarity calculation on any two keywords among all the keywords to obtain relevant similarity values; classifying all the keywords according to the relevant similarity values ​​to obtain a first phrase and a second phrase; performing cluster classification according to the first phrase using the text keywords corresponding to the first text to obtain a first classification result, and obtaining a first central word corresponding to each first sub-classification result and a first central weight corresponding to the first central word according to the first classification result; performing cluster classification according to the second phrase using the text keywords corresponding to the first text to obtain a first sub-classification result. A second classification result is obtained, and according to the second classification result, a second central word corresponding to each second sub-classification result and a second central weight corresponding to the second central word are obtained; according to the first central word and the second central word combined with the first central weight and the second central weight, a first associated word corresponding to the first central word in the second central word and a second associated word corresponding to the second central word in the first central word are determined; according to the first associated word and the second associated word, the first classification result and the second classification result are matched to obtain corresponding fusion similarities under multiple fusion phrases; according to the fusion similarities and the fusion phrases, the corresponding text similarities between the second texts are calculated; according to the text similarities, the second texts are text clustered to obtain the target clustering result.

[0073] Exemplarily, the second text is the first text processed with target keywords, and then keyword processing is performed on the second text according to the word frequency-inverse document frequency algorithm to identify the keywords therein to obtain text keywords corresponding to the second text, and all identified text keywords are aggregated together, repeated words are removed, and all keywords are formed.

[0074] Exemplarily, similarity calculation methods such as edit distance and cosine similarity are used to calculate the similarity between any two keywords among all keywords to obtain relevant similarity values, thereby classifying keywords with relevant similarity values ​​greater than a preset threshold into one category, and keywords with relevant similarity values ​​less than or equal to the preset threshold into another category, thereby forming a first phrase and a second phrase.

[0075] Exemplarily, based on the first phrase and combined with the text keywords corresponding to the second text, the similarity between the text keywords and the first phrase is calculated, and the second text is clustered according to the similarity result to obtain a first classification result. Similarly, based on the second phrase and combined with the text keywords corresponding to the second text, the similarity between the text keywords and the second phrase is calculated, and the second text is clustered according to the similarity result to obtain a second classification result.

[0076] Exemplarily, for each first sub-classification result in the first classification result, the first central word corresponding to the first sub-classification result is determined by calculating the comprehensive importance of the keywords in the sub-classification, for example, combining factors such as word frequency, position in the sub-classification, and the like, and a corresponding first central weight is assigned according to its importance. Similarly, for each second sub-classification result in the second classification result, the second central word corresponding to the second sub-classification result is determined by calculating the comprehensive importance of the keywords in the sub-classification, for example, combining factors such as word frequency, position in the sub-classification, and the like, and a corresponding second central weight is assigned according to its importance.

[0077] Exemplarily, the first central word and the second central word are compared, and the first central weight and the second central weight are combined to find the first associated word with a higher degree of association with the first central word in the second central word, and the second associated word corresponding to the second central word in the first central word.

[0078] Exemplarily, similarity is calculated based on the first associated words and the second associated words, so that the first sub-classification result in the first classification result and the second sub-classification result in the second classification result when the similarity is the largest are matched, so that the first central word corresponding to the first sub-classification result and the second central word corresponding to the second sub-classification result form a fused phrase, and then the first associated words and the second associated words corresponding to the formed fused phrase are deleted, and then the maximum value corresponding to the remaining similarity calculation result is obtained, so as to form another corresponding fused phrase, and then multiple fused phrases are obtained, and then for each fused phrase, the similarity between the included words and the corresponding weights are comprehensively considered to determine the corresponding fused similarity.

[0079] Exemplarily, the similarity between the semantic vectors of the fused phrase and the second text is calculated using a pre-trained language model, and this is used as the data association degree, and then the data association degree between each second text and each fused phrase is multiplied by the fusion similarity of the fused phrase, and then the calculation results of all fused phrases are summed. The result obtained is the text similarity corresponding to the second text under all fused phrases, and the calculated text similarity is used as input to cluster the second text using a clustering algorithm to finally obtain the target clustering result.

[0080] Specifically, the process of determining associated words and fused phrases helps to explore the deep connections between texts, not just the superficial lexical matching, but also discovers the potential semantic connections, providing more valuable information for further text clustering.

[0081] In some embodiments, the calculating the text similarity between the second texts according to the fused similarity and the fused phrase includes: determining a first subtext and a second subtext from the second text; obtaining a first frequency corresponding to each subkeyword in the fused phrase in the first subtext and a second frequency corresponding to each subkeyword in the second subtext; obtaining a third subclassification result corresponding to the fused phrase from the first classification result and a fourth subclassification result corresponding to the fused phrase from the second classification result; obtaining a first number corresponding to the keyword in the third subclassification result and a second number corresponding to the keyword in the fourth subclassification result; fusing the fused similarity corresponding to the fused phrase according to the first frequency, the second frequency, the first number and the second number to obtain the text similarity corresponding to the first subtext and the second subtext; wherein the text similarity is obtained according to the following formula:

[0082]

[0083] in, represents the text similarity between the i-th first subtext and the j-th second subtext, num2 represents the number of the fused phrases, num1 represents the number of words corresponding to the sub-keywords in the q-th fused phrase, represents the first frequency corresponding to the rth sub-keyword in the qth fused phrase in the ith first sub-text, represents the second frequency corresponding to the rth sub-keyword in the qth fused phrase in the jth second sub-text, represents the first number corresponding to the keywords in the third sub-classification result corresponding to the qth fused phrase, represents the second number corresponding to the keywords in the fourth sub-classification result corresponding to the qth fused phrase, represents the fusion similarity corresponding to the qth fused phrase.

[0084] Exemplarily, two subtexts are randomly selected from the second text and defined as the first subtext and the second subtext, respectively. The subtexts can be selected in sequence or randomly, and the specific method is determined according to actual needs.

[0085] Exemplarily, for each sub-keyword in the fused phrase, word frequency statistics are performed in the first sub-text and the second sub-text respectively. The first sub-text is scanned word by word, the number of times each sub-keyword appears is recorded, and the corresponding first frequency is obtained; similarly, the second sub-text is scanned, the number of times each sub-keyword appears is recorded, and the corresponding second frequency is obtained.

[0086] Exemplarily, a sub-classification result corresponding to the current fused phrase is found from the first classification result and determined as the third sub-classification result. Similarly, a sub-classification result corresponding to the fused phrase is found from the second classification result and determined as the fourth sub-classification result.

[0087] Exemplarily, the keywords in the third sub-category results are counted to obtain a first number; and the keywords in the fourth sub-category results are counted to obtain a second number.

[0088] Exemplarily, the first frequency, the second frequency, the first quantity, the second quantity, and the fusion similarity corresponding to the fused phrase obtained above are substituted into the following formula to calculate and obtain the text similarity:

[0089]

[0090] in, represents the text similarity between the i-th first subtext and the j-th second subtext, num2 represents the number of fused phrases, and num1 represents the number of words corresponding to the sub-keywords in the q-th fused phrase. represents the first frequency of the rth sub-keyword in the qth fused phrase in the i-th first sub-text, Indicates the second frequency of the rth sub-keyword in the qth fusion phrase in the jth second sub-text represents the first number of keywords in the third sub-classification result corresponding to the qth fused phrase, represents the second number of keywords in the fourth sub-classification result corresponding to the qth fused phrase, Represents the fusion similarity corresponding to the qth fused phrase.

[0091] Exemplarily, text similarity is calculated according to the above formula. After the algorithm performs deep mining on the first sub-text and the second sub-text, the extracted text information is contained in the fused phrase and the sub-classification results. Therefore, when calculating the text similarity between the first sub-text and the second sub-text, the ratio of the number of words in the fused phrase to the total number of text keywords is used as the weight to obtain the text similarity between the first sub-text and the second sub-text. The above method can reduce invalid calculations to a certain extent while considering the semantic information of words, and can better capture the similarity relationship between words.

[0092] Specifically, by analyzing fusion phrases and sub-keywords, and combining the information in different classification results, we can deeply explore the semantic information of the text. For example, the frequency of sub-keywords reflects its importance in the text, and the number of keywords in different classification results reflects the characteristics of the text under different classification systems. These are helpful to better understand the semantic connotation of the text, so that accurate text similarity provides a good foundation for subsequent text clustering, and can more accurately classify semantically similar texts into one category, improve the quality of clustering results, and make the clustered categories more representative and distinguishable.

[0093] In some embodiments, the performing similar text analysis on each sub-cluster in the target clustering result to obtain the common sentence of the sub-cluster and the second position of the common sentence includes: obtaining the corresponding third sub-text and fourth sub-text in the sub-cluster, and performing text segmentation on the third sub-text to obtain a first segmentation result and performing text segmentation on the fourth sub-text to obtain a second segmentation result; obtaining the third sub-sentence corresponding to the first segmentation result and the fourth sub-sentence corresponding to the second segmentation result; performing keyword recognition on the third sub-sentence to obtain a first keyword and performing keyword recognition on the fourth sub-sentence to obtain a second keyword; performing intersection processing on the first keyword and the second keyword to obtain the same keyword corresponding to the third sub-sentence and the fourth sentence and the target corresponding to the same keyword The method comprises the steps of: performing quantitative counting on the first keyword to obtain a third quantity, and performing logarithmic solution on the third quantity to obtain a first result; performing quantitative counting on the second keyword to obtain a fourth quantity, and performing logarithmic solution on the fourth quantity to obtain a second result; summing the first result and the second result to obtain a target result, and performing ratio calculation on the target quantity and the target result to obtain the sentence similarity corresponding to the third sub-sentence and the fourth sub-sentence; determining the similar sentences corresponding to the third sub-text and the fourth sub-text according to the sentence similarity; performing data statistics on the similar sentences according to the sub-class cluster to obtain the general sentence corresponding to the sub-class cluster; and performing position search in the sub-class cluster according to the general sentence to obtain the second position corresponding to the general sentence.

[0094] Exemplarily, two subtexts are randomly selected from each subclass cluster of the target clustering result to obtain the corresponding third subtext and fourth subtext, and then the third subtext is segmented into independent parts according to punctuation marks (period, exclamation mark, question mark, etc.), line breaks or specific separation marks to obtain the first segmentation result. The fourth subtext is segmented using the same method as the segmentation of the third subtext to obtain the second segmentation result.

[0095] Exemplarily, the third sub-sentence is extracted from the first segmentation result, and the fourth sub-sentence is extracted from the second segmentation result. The third sub-sentence and the fourth sub-sentence are usually short sentences with complete semantics, so the third sub-sentence is analyzed using keyword recognition such as a word frequency statistics-based, graph-based sorting algorithm to find representative and important words therein to obtain the first keyword, and the fourth sub-sentence is processed using the same keyword recognition method to obtain the second keyword.

[0096] Exemplarily, the first keyword and the second keyword are compared to find out the keywords they contain in common. These common keywords are the same keywords corresponding to the third sub-sentence and the fourth sub-sentence, and the number of the same keywords is counted to obtain the target number, and then the number of the first keyword is counted to obtain the third number, and then the logarithm of the third number is taken to obtain the first result. The number of the second keyword is counted to obtain the fourth number, and then the logarithm of the fourth number is taken to obtain the second result.

[0097] Exemplarily, the first result and the second result are added to obtain a target result, and the target number is divided by the target result to obtain the corresponding sentence similarity between the third sub-sentence and the fourth sub-sentence.

[0098] Exemplarily, a similarity threshold is set, and when the sentence similarity between the third sub-sentence and the fourth sub-sentence is greater than the threshold, the two sub-sentences are considered to be similar sentences. Thus, all third sub-sentences in the third sub-text and all fourth sub-sentences in the fourth sub-text are calculated pairwise to obtain similar sentences between the third sub-text and the fourth sub-text.

[0099] Exemplarily, data statistics are performed on all similar sentences in the entire subclass cluster, such as counting the frequency of occurrence of each similar sentence. Similar sentences with higher frequency of occurrence can be determined as common sentences corresponding to the subclass cluster.

[0100] Exemplarily, in all texts of the sub-class cluster, the positions where the common sentences appear are searched one by one, and the position information is recorded, so as to obtain the second position corresponding to the common sentence.

[0101] Specifically, through meticulous text processing and similarity calculation, we can accurately find the similarities between texts in sub-cluster and extract common sentences. These common sentences reflect the fixed sentences corresponding to the sub-cluster texts and are irrelevant to the core events that need to be conveyed in the official document content, thus providing good support for the subsequent generation of official document content with certain requirements or specifications.

[0102] Step S105: perform keyword recognition on the general sentence to obtain a second word and a second type of the second word.

[0103] Exemplarily, the naming recognition model is a bidirectional long short-term memory network and conditional random field model based on deep learning, so that the naming recognition model is used to perform sentence recognition on general sentences to obtain the word type identifier corresponding to each word in the general sentence, and determine the second word corresponding to the general sentence and the second type corresponding to the second word according to the word type identifier.

[0104] Step S106: extract rules from the key sentence according to the first word and the first type to obtain a first sentence rule.

[0105] Exemplarily, the key sentence is classified into event categories through a machine learning classification algorithm to obtain a first event type corresponding to the key sentence, and then the first event association between the first type and the first event type is determined based on the first word, so that the first word in the key sentence is replaced with the first type to obtain the first rule, and then the first sentence rule corresponding to the key sentence is determined based on the first event association between the first type and the first event type and the first rule.

[0106] For example, extract event-related features from the key sentence, such as the subject of the event (name, organization name, company name), action, time, place, etc. Therefore, replace the first word in the key sentence with the corresponding first type, and mark the first type with brackets or double quotes to distinguish it from other content. Then determine the relationship between the first type represented by the first word and the determined event type. For example, if the first word is "school" and the event type is "education event", then it can be analyzed that the type "school" is closely related to "education event", and determine the event relationship between the first type and the event type as the event subject. Therefore, replace the first word in the key sentence with the first type. For example, if the first word is "school" and the first type is "institution name", the key sentence "schools are prohibited from charging random fees" becomes "[institution name] is prohibited from charging random fees".

[0107] Step S107: extract rules from the general sentence according to the second word and the second type to obtain a second sentence rule.

[0108] Exemplarily, a general statement is classified into event categories through a machine learning classification algorithm to obtain a second event type corresponding to the general statement, and then the second event association between the second type and the second event type is determined based on the second word, so that the second word in the general statement is replaced with the second type to obtain the second rule, and then the second statement rule corresponding to the general statement is determined based on the first event association between the second type and the second event type and the second rule.

[0109] It should be noted that the common sentences are the commonly used sentences corresponding to the official document content when it is released, which are often unrelated to the event that needs to be released. The key sentences are the event information corresponding to the official document content when it is released.

[0110] Step S108: Obtain a target keyword of the event to be announced and a third type of the target keyword.

[0111] Exemplarily, an event to be published that a target user needs to publish is obtained, and then the event to be published is identified according to a named entity recognition model to obtain the word type corresponding to each word in the event to be published, and then the target keyword of the event to be published and the third type of the target keyword are obtained according to the word type.

[0112] Step S109: Generate target official document content of the event to be announced according to the target keyword and the third type in combination with the first sentence rule and the second sentence rule according to the first position and the second position.

[0113] Exemplarily, similarity calculation is performed between the event to be published and the event types corresponding to the first sentence rule and the second sentence rule, so as to filter out the first association rule and the second association rule corresponding to the event to be published from the first sentence rule and the second sentence rule.

[0114] Exemplarily, the event relationship between the target keyword and the event to be announced is determined, so that the target keyword is generated into a corresponding target key sentence according to the first association rule based on the event relationship in combination with the third type, and the target keyword is generated into a corresponding target general sentence according to the second association rule based on the event relationship in combination with the third type, so that the target official document content of the event to be announced is generated in the first position in the first text according to the first association rule and in the second position in the first text according to the second association rule.

[0115] In some embodiments, generating the target official document content of the event to be announced according to the first position and the second position based on the target keyword and the third type in combination with the first sentence rule and the second sentence rule includes: obtaining a first generated sentence according to the target keyword and the third type in combination with the first sentence rule; obtaining a second generated sentence according to the target keyword and the third type in combination with the second sentence rule; merging the first generated sentence and the second generated sentence according to the first position and the second position to obtain multiple initial official document contents; performing text quality assessment on the initial official document content according to a quality assessment model to obtain a target text quality corresponding to the initial official document content; and screening the target official document content of the event to be announced from the initial official document content according to the target text quality.

[0116] Exemplarily, similarity calculation is performed between the event to be published and the event types corresponding to the first sentence rule and the second sentence rule, so as to filter out the first association rule and the second association rule corresponding to the event to be published from the first sentence rule and the second sentence rule.

[0117] Exemplarily, the event relationship between the target keyword and the event to be announced is determined, so that the target keyword is generated into a corresponding first generated sentence according to the first association rule based on the event relationship in combination with the third type, and the target keyword is generated into a corresponding second generated sentence according to the second association rule based on the event relationship in combination with the third type.

[0118] Exemplarily, the first position and the second position indicate the arrangement order and combination method of the first generated sentence and the second generated sentence when they are merged, so that according to the requirements of the first position and the second position, the first generated sentence and the second generated sentence are merged to obtain the initial merged content, and then the text is expanded based on the initial merged content with the help of a text generation model to obtain multiple initial official document contents. The text generation model can be a model based on a neural network or a model based on deep learning.

[0119] For example, a model that can comprehensively evaluate the text quality is constructed. This model can consider multiple factors, such as grammatical correctness, semantic coherence, logical rationality, information integrity, etc. Then, each initial document content is input into the quality evaluation model, and it is scored according to the evaluation indicators and standards of the model to obtain the target text quality corresponding to each initial document content.

[0120] Exemplarily, a threshold of target text quality is set to filter out the document content with a quality greater than the threshold from all initial document content, and then the document content with a target text quality greater than the threshold is the target document content of the event to be announced.

[0121] See also Figure 2 , Figure 2An artificial intelligence-based document content generation device 200 is provided for an embodiment of the present application. The artificial intelligence-based document content generation device 200 includes a sentence recognition module 201, a first word recognition module 202, a replacement processing module 203, a cluster analysis module 204, a second word recognition module 205, a first rule extraction module 206, a second rule extraction module 207, a data acquisition module 208, and an official document generation module 209, wherein the sentence recognition module 201 is used to perform sentence recognition on a first text to obtain a key sentence and a first position of the key sentence; the first word recognition module 202 is used to perform keyword recognition on the key sentence to obtain a first word and a first type of the first word; the replacement processing module 203 is used to perform word replacement on the first word in the first text according to the first type to obtain a second text; the cluster analysis module 204 is used to perform text clustering on the second text to obtain a target cluster The target clustering result is obtained by performing similar text analysis on each sub-cluster in the target clustering result to obtain the common sentence of the sub-cluster and the second position of the common sentence; a second word recognition module 205 is used to perform keyword recognition on the common sentence to obtain the second word and the second type of the second word; a first rule extraction module 206 is used to perform rule extraction on the key sentence according to the first word and the first type to obtain the first sentence rule; a second rule extraction module 207 is used to perform rule extraction on the common sentence according to the second word and the second type to obtain the second sentence rule; a data acquisition module 208 is used to obtain the target keyword of the event to be announced and the third type of the target keyword; an official document generation module 209 is used to generate the target official document content of the event to be announced according to the first position and the second position in combination with the first sentence rule and the second sentence rule according to the target keyword and the third type.

[0122] In some implementations, the official document content generation device 200 based on artificial intelligence can be applied to a terminal device.

[0123] It should be noted that those skilled in the art can clearly understand that, for the sake of convenience and brevity of description, the specific working process of the official document content generation device 200 based on artificial intelligence described above can refer to the corresponding process in the aforementioned official document content generation method embodiment based on artificial intelligence, and will not be repeated here.

[0124] See also Figure 3 , Figure 3 A schematic block diagram of the structure of a terminal device provided in an embodiment of the present invention.

[0125] like Figure 3As shown, the terminal device 300 includes a processor 301 and a memory 302 , and the processor 301 and the memory 302 are connected via a bus 303 , such as an I2C (Inter-integrated Circuit) bus.

[0126] Specifically, the processor 301 is used to provide computing and control capabilities to support the operation of the entire terminal device. The processor 301 can be a central processing unit (CPU), and the processor 301 can also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0127] Specifically, the memory 302 may be a Flash chip, a read-only memory (ROM) disk, an optical disk, a USB flash drive, or a mobile hard disk.

[0128] Those skilled in the art will understand that Figure 3 The structure shown in the figure is only a block diagram of a partial structure related to the embodiment of the present invention, and does not constitute a limitation on the terminal device to which the embodiment of the present invention is applied. The specific server may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0129] The processor is used to run a computer program stored in the memory, and implement any one of the methods for generating official document content based on artificial intelligence provided by the embodiments of the present invention when executing the computer program.

[0130] In one embodiment, the processor is used to run a computer program stored in the memory, and implements the following steps when executing the computer program:

[0131] Performing sentence recognition on the first text to obtain a key sentence and a first position of the key sentence;

[0132] Performing keyword recognition on the key sentence to obtain a first word and a first type of the first word;

[0133] Performing word replacement on the first word in the first text according to the first type to obtain a second text;

[0134] Performing text clustering according to the second text to obtain a target clustering result, and performing similar text analysis on each sub-cluster in the target clustering result to obtain a common sentence of the sub-cluster and a second position of the common sentence;

[0135] Performing keyword recognition on the general sentence to obtain a second word and a second type of the second word;

[0136] Extracting rules from the key sentence according to the first word and the first type to obtain a first sentence rule;

[0137] Extracting rules from the general sentence according to the second word and the second type to obtain a second sentence rule;

[0138] Obtaining a target keyword of an event to be announced and a third type of the target keyword;

[0139] The target official document content of the event to be announced is generated according to the target keyword and the third type in combination with the first sentence rule and the second sentence rule according to the first position and the second position.

[0140] It should be noted that technical personnel in the relevant field can clearly understand that, for the convenience and brevity of description, the specific working process of the terminal device described above can refer to the corresponding process in the aforementioned embodiment of the document content generation method based on artificial intelligence, and will not be repeated here.

[0141] An embodiment of the present invention also provides a storage medium for computer-readable storage, wherein the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of any one of the methods for generating official document content based on artificial intelligence provided in the description of the embodiment of the present invention.

[0142] The storage medium may be an internal storage unit of the terminal device described in the foregoing embodiment, such as a hard disk or memory of the terminal device. The storage medium may also be an external storage device of the terminal device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), etc., equipped on the terminal device.

[0143] It will be appreciated by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In a hardware embodiment, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or transient medium). As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0144] It should be understood that the term "and / or" used in the present specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, including these combinations. It should be noted that, in this article, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "including a..." does not exclude the presence of other identical elements in the process, method, article or system including the element.

[0145] The serial numbers of the embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments. The above description is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.

Claims

1. A method for generating official document content based on artificial intelligence, characterized in that: The method comprises: Performing sentence recognition on the first text to obtain a key sentence and a first position of the key sentence; Performing keyword recognition on the key sentence to obtain a first word and a first type of the first word; Performing word replacement on the first word in the first text according to the first type to obtain a second text; Performing text clustering according to the second text to obtain a target clustering result, and performing similar text analysis on each sub-cluster in the target clustering result to obtain a common sentence of the sub-cluster and a second position of the common sentence; Performing keyword recognition on the general sentence to obtain a second word and a second type of the second word; Extracting rules from the key sentence according to the first word and the first type to obtain a first sentence rule; Extracting rules from the general sentence according to the second word and the second type to obtain a second sentence rule; Obtaining a target keyword of an event to be announced and a third type of the target keyword; The target official document content of the event to be announced is generated according to the target keyword and the third type in combination with the first sentence rule and the second sentence rule according to the first position and the second position.

2. The method according to claim 1, characterized in that The step of performing sentence recognition on the first text to obtain a key sentence and a first position of the key sentence includes: Perform keyword extraction on the first text to obtain a third word corresponding to the first text and a word weight corresponding to the third word; Performing sentence segmentation on the first text to obtain an initial sentence, and determining a sentence weight corresponding to the initial sentence according to the third word and the word weight; Determine any one of the initial sentences as a first sentence, and obtain a second sentence corresponding to the first sentence in an adjacent state from the initial sentence; Obtaining a first weight corresponding to the first sentence and a second weight corresponding to the second sentence from the sentence weights; Adjusting the first weight according to the second weight to obtain a target weight corresponding to the first sentence; The initial sentence is subjected to sentence recognition according to the target weight to obtain the key sentence, and the first position corresponding to the key sentence is obtained from the first text according to the key sentence.

3. The method according to claim 2, characterized in that The step of weight-adjusting the first weight according to the second weight to obtain a target weight corresponding to the first sentence includes: Obtaining a first sequence of statements corresponding to a left position of the second statement from the initial statement, and obtaining a second sequence of statements corresponding to a right position of the second statement from the initial statement; Obtaining the associated keywords corresponding to each first sub-sentence in the first sequence of sentences from the third words, and obtaining the relevant weights corresponding to the associated keywords from the word weights; Obtaining a third weight corresponding to each of the first sub-sentences in the first sequence of sentences from the sentence weights and obtaining a fourth weight corresponding to each of the second sub-sentences in the second sequence of sentences from the sentence weights; According to the relevant weight, combined with the third weight and the fourth weight, the first weight is weight-adjusted using the second weight to obtain the target weight corresponding to the first sentence; The target weight is obtained according to the following formula: in, represents the target weight corresponding to the i-th first sentence, represents the adjustment parameter, represents the first weight corresponding to the i-th first sentence, n represents the number of sentences corresponding to the first sequence of sentences, y represents the number of associated keywords corresponding to the h-th first sub-sentence, represents the relevant weight corresponding to the kth associated keyword corresponding to the hth first sub-sentence, represents the third weight corresponding to the hth first sub-sentence, g represents the number of sentences corresponding to the second sequence of sentences, represents the fourth weight corresponding to the t-th second sub-statement, Represents the second weight of the second sentence corresponding to the i-th adjacent state of the first sentence.

4. The method according to claim 1, characterized in that: The performing text clustering according to the second text to obtain a target clustering result includes: Performing keyword recognition on the second text to obtain text keywords corresponding to the second text; Merging the text keywords to obtain all the keywords, and calculating the similarity between any two keywords in all the keywords to obtain relevant similarity values; Classifying all the keywords according to the relevant similarity values ​​to obtain a first word group and a second word group; Performing cluster classification using the text keywords corresponding to the second text according to the first phrase to obtain a first classification result, and obtaining a first central word corresponding to each first sub-classification result and a first central weight corresponding to the first central word according to the first classification result; Performing cluster classification using the text keywords corresponding to the second text according to the second phrase to obtain a second classification result, and obtaining a second central word corresponding to each second sub-classification result and a second central weight corresponding to the second central word according to the second classification result; Determine, according to the first central word and the second central word in combination with the first central weight and the second central weight, a first associated word corresponding to the first central word in the second central word and a second associated word corresponding to the second central word in the first central word; Matching the first classification result and the second classification result according to the first associated words and the second associated words to obtain corresponding fusion similarities under multiple fusion phrases; Calculating the text similarity corresponding to the second texts according to the fusion similarity and the fusion phrase; The second text is clustered according to the text similarity to obtain the target clustering result.

5. The method according to claim 4, characterized in that The calculating the text similarity corresponding to the second texts according to the fusion similarity and the fusion phrases includes: determining a first subtext and a second subtext from the second text; Obtaining a first frequency corresponding to each sub-keyword in the fused phrase in the first sub-text and a second frequency corresponding to each sub-keyword in the second sub-text; Obtain a third sub-classification result corresponding to the fused phrase from the first classification result and obtain a fourth sub-classification result corresponding to the fused phrase from the second classification result; Obtaining a first quantity corresponding to the keyword in the third sub-category result and obtaining a second quantity corresponding to the keyword in the fourth sub-category result; According to the first frequency, the second frequency, the first number, and the second number, the fusion similarity corresponding to the fusion phrase is fused to obtain the text similarity corresponding to the first subtext and the second subtext; The text similarity is obtained according to the following formula: in, represents the text similarity between the i-th first subtext and the j-th second subtext, num2 represents the number of the fused phrases, num1 represents the number of words corresponding to the sub-keywords in the q-th fused phrase, represents the first frequency corresponding to the rth sub-keyword in the qth fused phrase in the ith first sub-text, represents the second frequency corresponding to the rth sub-keyword in the qth fused phrase in the jth second sub-text, represents the first number corresponding to the keywords in the third sub-classification result corresponding to the qth fused phrase, represents the second number corresponding to the keywords in the fourth sub-classification result corresponding to the qth fused phrase, represents the fusion similarity corresponding to the qth fused phrase.

6. The method according to claim 1, characterized in that The performing similar text analysis on each sub-cluster in the target clustering result to obtain a common sentence of the sub-cluster and a second position of the common sentence includes: Obtaining a third subtext and a fourth subtext corresponding to the subclass cluster, and performing text segmentation on the third subtext to obtain a first segmentation result and performing text segmentation on the fourth subtext to obtain a second segmentation result; Obtaining a third sub-sentence corresponding to the first segmentation result and a fourth sub-sentence corresponding to the second segmentation result; Performing keyword recognition on the third sub-sentence to obtain a first keyword and performing keyword recognition on the fourth sub-sentence to obtain a second keyword; Performing intersection processing on the first keyword and the second keyword to obtain the same keyword corresponding to the third sub-sentence and the fourth sub-sentence and the target number corresponding to the same keyword; Counting the first keyword to obtain a third quantity, and performing logarithmic solution on the third quantity to obtain a first result; Counting the number of the second keyword to obtain a fourth number, and performing logarithmic solution on the fourth number to obtain a second result; Summing the first result and the second result to obtain a target result, and performing a ratio calculation between the target quantity and the target result to obtain a sentence similarity corresponding to the third sub-sentence and the fourth sub-sentence; Determining similar sentences corresponding to the third subtext and the fourth subtext according to the sentence similarity; Performing data statistics on the similar sentences according to the sub-class cluster to obtain the common sentences corresponding to the sub-class cluster; A position search is performed in the subclass cluster according to the general statement to obtain the second position corresponding to the general statement.

7. The method according to claim 1, characterized in that The step of generating the target official document content of the event to be announced according to the target keyword and the third type in combination with the first sentence rule and the second sentence rule according to the first position and the second position includes: Obtaining a first generated sentence according to the target keyword and the third type in combination with the first sentence rule; Obtaining a second generated sentence according to the target keyword and the third type combined with the second sentence rule; Merging the first generated sentence and the second generated sentence according to the first position and the second position to obtain a plurality of initial official document contents; Performing text quality assessment on the initial official document content according to the quality assessment model to obtain a target text quality corresponding to the initial official document content; The target official document content of the event to be announced is obtained by screening the initial official document content according to the target text quality.

8. An official document content generation device based on artificial intelligence, characterized in that: include: A sentence recognition module, used for performing sentence recognition on the first text to obtain a key sentence and a first position of the key sentence; A first word recognition module, configured to perform keyword recognition on the key sentence to obtain a first word and a first type of the first word; a replacement processing module, configured to replace the first word in the first text according to the first type to obtain a second text; A cluster analysis module, configured to perform text clustering according to the second text to obtain a target clustering result, and perform similar text analysis on each sub-cluster in the target clustering result to obtain a common sentence of the sub-cluster and a second position of the common sentence; A second word recognition module, used for performing keyword recognition on the general sentence to obtain a second word and a second type of the second word; A first rule extraction module, configured to extract rules from the key sentence according to the first word and the first type to obtain a first sentence rule; A second rule extraction module, configured to extract rules from the general sentence according to the second word and the second type to obtain a second sentence rule; A data acquisition module, used to obtain a target keyword of an event to be announced and a third type of the target keyword; The official document generation module is used to generate the target official document content of the event to be announced according to the first position and the second position in combination with the first sentence rule and the second sentence rule according to the target keyword and the third type.

9. A terminal device, characterized in that: The terminal device includes a processor and a memory; The memory is used to store computer programs; The processor is used to execute the computer program and implement the method for generating official document content based on artificial intelligence as described in any one of claims 1 to 7 when executing the computer program.

10. A computer storage medium for computer storage, characterized in that: The computer storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the document content generation method based on artificial intelligence as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Text classification method and device, electronic equipment and readable storage medium

    CN112597312A

  • Text intention classification method and device, equipment and storage medium

    CN114860942A

  • Sentence generation method for expanding corpus, electronic equipment and storage medium

    CN115169359A

Cited By

  • Automatic official document generation method based on dynamic rules

    CN120975050A