Text generation method and related equipment

Through semantic analysis and feedback iteration methods, combined with termbases and text generation factors, the problem of insufficient logic and integrity of document generation in the existing technology is solved, and efficient and professional text generation is achieved.

CN120387457APending Publication Date: 2025-07-29NEW PRIME NUMBER (BEIJING) DATA CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510162528.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing technical document writing tools cannot deeply understand the semantic structure of documents, resulting in insufficient logic and integrity of the generated content, affecting the efficiency and quality of document generation.

Method used

By obtaining the user input text and term database, semantic analysis is performed to determine the keyword list, and using the text generation model to combine the text generation factor to generate preliminary output text, iteratively update the keyword list and factor through feedback information, and finally generate output text that conforms to professionalism and logic.

Benefits of technology

It significantly improves the professionalism and logic of the generated content, shortens the time for writing technical documents, and improves text generation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387457A_ABST
    Figure CN120387457A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a text generation method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring an input text input by a user and a term library; determining a keyword list through semantic analysis; determining a text generation factor; generating a first output text through a text generation model according to the term library, the keyword list and the text generation factor; the feedback information is analyzed and processed, the keyword list and the text generation factor are updated, and a second output text is generated according to the updated keyword list and text generation factor. The method is combined with the corresponding term library, so that the specialty of the generated content is greatly improved. Moreover, the feedback information is analyzed through multiple rounds of iteration, and then the keyword list and the text generation factor in the text generation process are adjusted, so that the text generation efficiency can be greatly improved, and the writing time of the technical document is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a text generation method, apparatus, electronic device, computer-readable medium, and computer program product. Background Art

[0002] Currently, the writing of technical documents mainly relies on manual work. Most existing document writing tools mainly focus on content editing and format typesetting, and fail to deeply understand the semantic structure of the documents and provide dynamic optimization suggestions. For example, in the case where the user input information is incomplete or the document logic is incoherent, the existing tools cannot identify these problems and give effective guidance and supplementary suggestions, resulting in the user having to repeatedly modify the document, which is time-consuming and laborious.

[0003] Although large language models (LLMs) can understand and generate high-quality texts and are widely used in natural language processing tasks. However, the existing large language models are difficult to adapt to the requirements in different fields and complex scenarios, and cannot ensure that the generated content is logical and complete. This limitation not only reduces the document generation efficiency, but also affects the quality of the final document, making the process of writing technical documents long and inefficient. Summary of the Invention

[0004] Embodiments of the present disclosure provide a text generation method, apparatus, electronic device, computer-readable medium, and computer program product.

[0005] According to a first aspect of the embodiments of the present disclosure, a text generation method is provided, including: obtaining an input text input by a user and a term library corresponding to the input text; determining a keyword list by performing semantic analysis on the input text; the keyword list includes at least one keyword; determining at least one text generation factor; the text generation factor is used to adjust text generation parameters; generating a first output text by a pre-trained text generation model according to the term library, the keyword list, and at least one text generation factor; obtaining feedback information on the first output text; updating the keyword list and at least one text generation factor by analyzing and processing the feedback information to obtain an updated keyword list and at least one text generation factor; generating a second output text by the text generation model according to the term library, the updated keyword list, and at least one text generation factor.

[0006] In some exemplary embodiments of the present disclosure, determining the keyword list by performing semantic analysis on the input text includes: performing word frequency analysis on the input text to screen out at least one high-frequency word; calculating the semantic similarity between each of the high-frequency words and the input text through a semantic analysis model; determining the word weight value corresponding to each of the high-frequency words; sorting the high-frequency words among the at least one high-frequency word according to the semantic similarity and the word weight value; and determining the keyword list according to the sorting of the high-frequency words.

[0007] In some exemplary embodiments of the present disclosure, the at least one text generation factor includes at least one of the following text generation factors: a temperature factor for adjusting the randomness of the generated content; a repetition penalty factor for adjusting the generation probability of repeated content; and a semantic adaptation factor for adjusting the matching degree between the generated content and the term library.

[0008] In some exemplary embodiments of the present disclosure, obtaining the feedback information on the first output text; and updating the keyword list and the at least one text generation factor by analyzing and processing the feedback information to obtain the updated keyword list and the at least one text generation factor, includes: obtaining the feedback text on the first output text; determining the feedback semantic information by performing semantic analysis on the feedback text; and updating the keyword list and the at least one text generation factor according to the feedback semantic information to obtain the updated keyword list and the at least one text generation factor.

[0009] In some exemplary embodiments of the present disclosure, obtaining the feedback information on the first output text; and updating the keyword list and the at least one text generation factor by analyzing and processing the feedback information to obtain the updated keyword list and the at least one text generation factor, includes: obtaining the feedback text on the first output text; determining the corresponding feedback type by performing type analysis on the feedback text through a feedback classification model; and updating the keyword list and the at least one text generation factor according to the feedback text by a feedback processing module corresponding to the feedback type to obtain the updated keyword list and the at least one text generation factor.

[0010] In some exemplary embodiments of the present disclosure, obtaining feedback information on the first output text; updating the keyword list and at least one text generation factor by analyzing and processing the feedback information to obtain an updated keyword list and at least one text generation factor, including: performing text proofreading processing on the first output text through a proofreading model to obtain text proofreading information; updating the keyword list and at least one text generation factor according to the text proofreading information to obtain an updated keyword list and at least one text generation factor.

[0011] In some exemplary embodiments of the present disclosure, obtaining feedback information on the first output text; updating the keyword list and at least one text generation factor by analyzing and processing the feedback information to obtain an updated keyword list and at least one text generation factor, including: obtaining the feedback text on the first output text; obtaining at least one feedback problem by analyzing and processing the feedback text; determining the feedback priority corresponding to each feedback problem; and updating the keyword list and at least one text generation factor by analyzing and processing the corresponding feedback problem in sequence according to the feedback priority to obtain an updated keyword list and at least one text generation factor.

[0012] In some exemplary embodiments of the present disclosure, obtaining the input text input by the user and the term library corresponding to the input text further includes: determining a matching text template according to the input text; at least one information element corresponds to the text template; determining the number of information elements covered by the input text according to the input text and at least one information element; determining an information integrity evaluation value of the input text according to the number of covered information elements; and issuing a guiding question in response to the information integrity evaluation value not reaching a preset integrity threshold; the guiding question is used to guide the user to supplement the uncovered information elements.

[0013] In some exemplary embodiments of the present disclosure, issuing a guiding question in response to the information integrity evaluation value not reaching a preset integrity threshold includes: in response to the information integrity evaluation value not reaching a preset integrity threshold, determining at least one guiding question generation factor according to the information integrity evaluation value; and generating and issuing the guiding question by a guiding question generation model according to the uncovered information elements and at least one guiding question generation factor.

[0014] According to a second aspect of the embodiments of the present disclosure, there is provided a text generation device, including: an input text acquisition unit, configured to acquire an input text input by a user and a term library corresponding to the input text; a keyword determination unit, configured to determine a keyword list by performing semantic analysis on the input text; the keyword list includes at least one keyword; a text generation factor determination unit, configured to determine at least one text generation factor; the text generation factor is used to adjust text generation parameters; a first text generation unit, configured to generate a first output text according to the term library, the keyword list and at least one text generation factor through a pre-trained text generation model; a feedback information unit, configured to acquire feedback information on the first output text; by analyzing and processing the feedback information, update the keyword list and at least one text generation factor to obtain an updated keyword list and at least one text generation factor; a second text generation unit, configured to generate a second output text according to the term library, the updated keyword list and at least one text generation factor through the text generation model.

[0015] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, characterized by including: a processor; a memory for storing executable instructions executable by the processor; wherein, the processor is configured to execute the executable instructions to implement any one of the text generation methods.

[0016] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute any one of the text generation methods.

[0017] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program, and the computer program is executed to implement any one of the text generation methods by a processor.

[0018] The text generation method provided by the embodiments of the present disclosure acquires an input text input by a user and a term library; determines a keyword list through semantic analysis; determines text generation factors; generates a first output text through a text generation model according to the term library, the keyword list and the text generation factors; updates the keyword list and the text generation factors by analyzing and processing the feedback information, and generates a second output text according to the updated keyword list and text generation factors. By combining with the corresponding term library, this method greatly improves the professionalism of the generated content. Moreover, by iteratively analyzing the feedback information in multiple rounds and then adjusting the keyword list and text generation factors in the text generation process, the text generation efficiency can be greatly improved and the writing time of technical documents can be shortened.

[0019] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0021] Figure 1 is a flowchart of a text generation method shown according to an exemplary embodiment.

[0022] Figure 2 is a flowchart of a keyword list determination process shown according to an example.

[0023] Figure 3 is a flowchart of a feedback information processing process shown according to an example Figure 1 .

[0024] Figure 4 is a flowchart of a feedback information processing process shown according to an example Figure 2 .

[0025] Figure 5 is a flowchart of a feedback information processing process shown according to an example Figure 3 .

[0026] Figure 6 is a flowchart of a feedback information processing process shown according to an example Figure 4 .

[0027] Figure 7 is a flowchart of an input text supplementing process shown according to an example.

[0028] Figure 8 is a block diagram of a text generation device shown according to an exemplary embodiment. DETAILED DESCRIPTION

[0029] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. Like reference numerals in the figures denote like or similar parts, and thus their repeated description will be omitted.

[0030] The features, structures, or characteristics described in this disclosure may be combined in one or more embodiments in any suitable manner. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of this disclosure. However, those skilled in the art will realize that one or more of the specific details may be omitted to practice the technical solutions of this disclosure, or other methods, components, devices, steps, etc. may be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this disclosure.

[0031] The accompanying drawings are only schematic illustrations of this disclosure, and the same reference numerals in the drawings represent the same or similar parts, so repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software form, or in at least one hardware module or integrated circuit, or in different networks and / or processor devices and / or microcontroller devices.

[0032] The flowcharts shown in the accompanying drawings are only illustrative, and do not necessarily include all the contents and steps, nor do they necessarily need to be executed in the described order. For example, some steps can be decomposed, while some steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0033] In this specification, the terms "a", "an", "the", "said", and "at least one" are used to indicate the existence of at least one element / component / etc.; the terms "comprising", "including", and "having" are used to mean an open inclusion and refer to the existence of additional elements / components / etc. in addition to the listed elements / components / etc.; the terms "first", "second", "third", etc. are only used as labels and are not a limitation on the quantity of their objects.

[0034] Next, each step of the method in the exemplary embodiments of this disclosure will be described in more detail with reference to the accompanying drawings and embodiments.

[0035] Figure 1 is a flowchart of a text generation method shown according to an exemplary embodiment. Figure 1 The method provided by the embodiment can be executed by any electronic device, such as a terminal device, a server, or jointly executed by a terminal device and a server, but this disclosure does not limit this.

[0036] In step S110, the input text input by the user and the term library corresponding to the input text are obtained.

[0037] In an embodiment of the present disclosure, a method for generating a technical document based on a large language model for an input text provided by a user is provided. First, the input text provided by the user is obtained. The input text can be a text generation prompt word provided by the user, an existing technical document, a technical literature link, or a meeting minutes. Of course, the input text can also be input in the form of voice or picture and converted into a corresponding text format through existing conversion technologies such as a voice recognition tool or an OCR tool.

[0038] In an embodiment of the present disclosure, a term library corresponding to the input text is further obtained. The term library is a professional term library corresponding to the technical field involved in the input text. By using the term library corresponding to the input text, the professionalism of the words in the input text can be corrected during the text generation process, and at the same time, the importance of each word in the input text in the professional technical field can be determined, so as to generate a text that better conforms to the professional technical field and improve the quality of the generated technical document.

[0039] In an exemplary embodiment, the text generation system can provide term libraries of multiple different technical fields at the same time. In step S110, the term library can be manually selected by the user to match the term library, or automatically matched to the corresponding term library after semantic analysis of the input text.

[0040] In step S120, a keyword list is determined by performing semantic analysis on the input text; the keyword list includes at least one keyword.

[0041] In an embodiment of the present disclosure, several keywords in the input text are identified and judged by performing semantic analysis on the input text. Based on the several keywords, a keyword list is formed. The keyword list filters out the non-key information in the input text and extracts the key information related to the user text generation for subsequent text generation.

[0042] In an exemplary embodiment, the process of determining a keyword list by semantic analysis of the input text mainly relies on natural language processing models and algorithms. First, the input text is tokenized to divide it into several words. Secondly, each word is converted into a corresponding word feature vector through methods such as convolutional processing. Then, a pre-trained language model such as BERT is called to perform in-depth semantic analysis on the text to generate context embedding vectors for each word. These vectors can capture the semantic information of words in a specific context and provide a basis for subsequent keyword extraction. Then, a semantic similarity algorithm (such as cosine similarity or Euclidean distance) is used to calculate the correlation between the feature vectors of each word and the overall feature vector of the input text. Through dynamic weight adjustment, the system can dynamically optimize the weight values of each word according to the context. Several keywords that can reflect the core theme of the input text are selected from each word based on the above correlation and weight values, and a keyword list is formed based on the several keywords. This method effectively solves the problem of insufficient context adaptability of traditional keyword extraction techniques, improves the accuracy and domain adaptability of keywords, and avoids the risk of the generated content deviating from the theme.

[0043] In step S130, at least one text generation factor is determined; the text generation factor is used to adjust text generation parameters.

[0044] In the embodiments of the present disclosure, at least one text generation factor related to text generation is determined. The text generation factor is used to adjust text generation parameters. Different text generation factors correspond to adjusting text generation parameters in different dimensions.

[0045] In an exemplary embodiment, the text generation factor can be manually set by the user, its initial value can be defaulted by the system, or the corresponding text generation factor can be determined based on the analysis of the input text. The present disclosure does not limit the specific determination method of the text generation factor.

[0046] In an exemplary embodiment, the at least one text generation factor may include at least one of the following text generation factors:

[0047] Temperature factor (T), which is used to adjust the randomness of the generated content. By this temperature factor, the diversity and certainty of the generated content are controlled, and the value range is usually [0.2, 1.0]. The smaller the temperature factor, the higher the certainty of the generated text content compared to the input text. The larger the temperature factor, the more diverse the generated text content will be. The temperature factor can affect text generation through the following formula:

[0048]

[0049] where P(w) represents the generation probability of ω, and E(w) represents the score of the corresponding word.

[0050] The repetition penalty factor (Rp) is used to adjust the generation probability of repeated content. By adjusting the generation probability of words with this repetition penalty factor, the generation of repeated content is reduced, and its value range is usually [1.1, 2.0]. This repetition penalty factor can affect text generation through the following formula:

[0051]

[0052] where P(w) represents the generation probability of ω, c(w) is the number of occurrences of word w in the generated sequence, and R p > 1 can effectively penalize the generation of repeated content.

[0053] The semantic adaptation factor (S α ) is used to adjust the matching degree between the generated content and the said term library. The adaptation degree between the generated content and the domain term library is measured by this semantic adaptation factor, and its value range is [0, 1]. By using this semantic adaptation factor, the professionalism of the generated content can be improved in combination with the above-mentioned term library. Let the weight of term t be W t , which is used to represent the importance weight of term t in the domain term library. This weight can be dynamically set to adapt to different generation requirements. The semantic adaptation factor of this generation process is calculated by the following formula:

[0054]

[0055] where S(w, t) represents the semantic similarity between word ω and term t, which can be calculated by the cosine similarity of word vectors of the pre-trained model. The larger the value, the higher the semantic similarity.

[0056] It should be noted that the above-provided text generation factors are only used for illustrative purposes and do not limit the protection scope of the present disclosure. Those skilled in the art can select one or several of these text generation factors according to actual needs, or design and adjust text generation factors in other dimensions, which should all be regarded as within the protection scope of the present disclosure.

[0057] In step S140, a first output text is generated by a pre-trained text generation model according to the term library, the keyword list, and at least one text generation factor.

[0058] In the related art, large language models (LLMs) can understand and generate high-quality text and are widely used in natural language processing tasks. Currently, text generation technologies based on large language models have been widely applied in fields such as dialogue systems and content creation, and the generation quality has been significantly improved. However, existing large language models are difficult to adapt to the requirements in different fields and complex scenarios and cannot ensure the logic and integrity of the generated content.

[0059] In the embodiments of the present disclosure, for the scenario of generating professional technical documents of the present disclosure based on a large language model, a text generation model is obtained through pre-training. The term library obtained in the foregoing steps, the keyword list extracted based on the input text, and at least one determined text generation factor are input into the text generation model. The text generation model generates a first output text according to the term library, the keyword list, and at least one text generation factor.

[0060] In the embodiments of the present disclosure, during the text generation process, the text generation model first introduces a domain term library, integrates professional vocabulary and knowledge in a specific domain into the model, and ensures that the generated content meets professional standards in terms of domain adaptability. Combining with the keyword list, the model can accurately capture the core concepts and themes of the generation target, providing a clear direction for the generation process. At the same time, the text generation factors dynamically adjust the generation parameters, optimize logical coherence and language fluency, so that the generated content not only conforms to domain specifications but also has the expression effect of natural language.

[0061] In the decoding stage, the Beam Search technology is adopted to generate the optimal sequence through multi-path search, ensuring the logical coherence and language fluency of the generated text. In addition, the generation of each text generation factor is dynamically adjusted to control the diversity of the generated content while avoiding repetition and redundancy, so that the generated content is both innovative and meets domain requirements. By combining the pre-trained model with a customized term library, the model can generate content for a specific technical field, significantly improving the logic and professionalism of the generated content.

[0062] Finally, the text generation model gradually generates a first output text that conforms to domain specifications, is logically clear, has fluent language, and has a certain degree of innovation under the comprehensive guidance of the term library, the keyword list, and the text generation factors.

[0063] In an exemplary embodiment, the user inputs the text: "Optimization of UAV power system". The system combines the domain term library to generate the first output text: "The present invention relates to an optimization scheme for the UAV power system, especially the dynamic balance mechanism between motor energy consumption and battery management." The system further optimizes the logical coherence and professionalism of the content by dynamically adjusting the generation temperature factor and semantic adaptation factor.

[0064] In step S150, feedback information on the first output text is obtained; by analyzing and processing the feedback information, the keyword list and at least one text generation factor are updated to obtain an updated keyword list and at least one text generation factor.

[0065] Since it is difficult for the first output text generated based on the user input text to directly achieve a satisfactory effect, the present disclosure provides a mechanism for regenerating the output text based on feedback information, dynamically adjusting the generated content through multiple rounds of optimization iterations, so as to generate an output text that better meets the user's expectations.

[0066] In an embodiment of the present disclosure, feedback information for the first output text is obtained. The feedback information may be further feedback text provided by the user for the first output text, or text calibration information generated by a pre-trained calibration model through calibration processing of the first output text.

[0067] In an embodiment of the present disclosure, by analyzing and processing the feedback information, the previously determined keyword list and / or at least one text generation factor are updated, so as to obtain an updated keyword list and at least one text generation factor. It should be noted that according to the content reflected by the feedback information, partial information in the keyword list and at least one text generation factor may be updated. In addition, according to actual needs, the correspondence between the feedback information and the updated content may be adjusted, and the present disclosure does not limit the mathematical relationship between the feedback information and the updated content.

[0068] In step S160, a second output text is generated by the text generation model according to the term library, the updated keyword list, and at least one text generation factor.

[0069] In an embodiment of the present disclosure, through the same text generation model in the aforementioned step S140, text generation is performed again according to the term library, the updated keyword list, and at least one text generation factor to obtain a second output text. The generation process of the specific output text has been introduced in the aforementioned step S140 and will not be repeated here.

[0070] It should be noted that the above steps S150 and S160 are a cyclic iteration process. The user can provide relevant feedback information based on the output text generated in the previous round, continuously update the keyword list and text generation factors through analysis and processing of the relevant feedback information, so as to regenerate an updated output text based on the user feedback information. Through multiple rounds of optimization iterations, the generated content is dynamically adjusted until an output text that meets the user's expectations is generated.

[0071] The text generation method provided by the embodiments of the present disclosure obtains the input text entered by the user and a thesaurus; determines a keyword list through semantic analysis; determines text generation factors; generates a first output text through a text generation model according to the thesaurus, the keyword list, and the text generation factors; analyzes and processes the feedback information, updates the keyword list and the text generation factors, and generates a second output text according to the updated keyword list and text generation factors. By combining with the corresponding thesaurus, this method greatly improves the professionalism of the generated content. Moreover, by iteratively analyzing the feedback information in multiple rounds and then adjusting the keyword list and the text generation factors in the text generation process, the text generation efficiency can be greatly improved, and the writing time of technical documents can be shortened.

[0072] Figure 2 is a flowchart showing the process of determining a keyword list according to an example. As Figure 2 shown, in the embodiments of the present disclosure, the process of determining the keyword list in the foregoing step S120 may include the following steps.

[0073] In step S210, perform a word frequency analysis on the input text to screen out at least one high-frequency word.

[0074] In the embodiments of the present disclosure, first, perform a word segmentation process on the input text to divide the input text into several words. Through the word frequency analysis of each word in the input text, at least one high-frequency word is screened out based on the word frequency of the word in the input text as a candidate keyword.

[0075] In an exemplary embodiment, the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm may be used to perform a word frequency analysis on the user input text. The TF-IDF algorithm is a statistical method for evaluating the importance of a word in a document. Its core idea is that the higher the frequency of a word in a document (TF), and the lower the frequency of the word in the entire corpus (IDF), the more representative the word is for the document. In the process of screening keywords based on word frequency analysis, first calculate the TF value of each word in the user input text to measure its occurrence frequency in the current document; then calculate the IDF value to measure the rarity of the word in the entire corpus; finally, multiply TF by IDF to obtain the TF-IDF value, and screen out the words with higher TF-IDF values as preliminary keywords. This method can effectively identify the words that are of important significance to the content of the document, and at the same time filter out the common but meaningless words.

[0076] In step S220, through a semantic analysis model, calculate the semantic similarity between each of the high-frequency words and the input text.

[0077] In the embodiments of the present disclosure, the semantic similarity between each of the high-frequency words and the overall context of the input text is calculated through a pre-trained semantic analysis model. This semantic similarity characterizes the relevance of the high-frequency word in the overall context. Currently, there are various semantic models that can perform the evaluation and calculation of the relevant similarity, and the present disclosure does not limit the specific model type of the semantic analysis model.

[0078] In an exemplary embodiment, the BERT model can be used as the semantic analysis model. BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language model based on the Transformer architecture, which captures the deep semantic information of the text through bidirectional context understanding. By calling pre-trained language models such as BERT to perform in-depth semantic analysis on the text, the context embedding vectors of each word are generated. These vectors can capture the semantic information of the word in a specific context and provide a basis for subsequent keyword extraction. Then, the semantic similarity algorithm (such as cosine similarity or Euclidean distance) is used to calculate the relevance between the feature vectors of each high-frequency word and the overall feature vector of the input text, and the semantic similarity is obtained.

[0079] In step S230, the word weight values corresponding to each of the high-frequency words are determined.

[0080] In the embodiments of the present disclosure, each high-frequency word is also evaluated from the perspective of the word weight value. The word weight values corresponding to each high-frequency word are determined. The word weight value is used to characterize the importance of the corresponding word from different dimensions.

[0081] In an exemplary embodiment, the word weight value may include weight values of multiple different dimensions. For example, it may include any one or more of the following word weight values.

[0082] Word frequency weight value (W f ), which is used to measure the frequency of occurrence of the keyword in the user input. It can be expressed by the following formula:

[0083]

[0084] where f k represents the number of occurrences of the keyword k in the current input text. F represents the total number of words in the input text.

[0085] Domain relevance weight value (W u ), which is used to measure the semantic similarity between the keyword and the term library. It can be expressed by the following formula:

[0086] W d =S(k, D)

[0087] Among them, S(k, D) represents the semantic similarity between keyword k and the thesaurus D, and its value range is [0, 1].

[0088] The user's historical preference weight value (W u ), which is used to measure the preference degree of the keyword in the user's historical input. It can be expressed by the following formula:

[0089]

[0090] Among them, c k represents the number of times keyword k is used by the user in the historical input. C represents the total number of words in the historical input. Here, "historical" refers to the text content input by the user during the process of generating past texts.

[0091] Based on the word weight values of the above various different dimensions, the final word weight value of the high-frequency word can be calculated by the following formula:

[0092] W fina1 = α·W f + β·W d + γ·W u

[0093] Among them, α, β, γ: dynamic adjustment coefficients, which are used to balance the influences of word frequency, semantic similarity, and user preference. Their values are dynamically adjusted by the system according to different scenarios.

[0094] It should be noted that the word weight values provided above are only used for illustrative purposes and do not limit the protection scope of the present disclosure. Those skilled in the art can select one or several of the word weight values according to actual needs, or design word weight values based on other dimensions, which should all be regarded as within the protection scope of the present disclosure.

[0095] In step S240, according to the semantic similarity and the word weight value, the high-frequency words among the at least one high-frequency word are sorted.

[0096] In the embodiment of the present disclosure, the semantic similarity and the word weight value corresponding to each high-frequency word are obtained according to the foregoing steps. Based on the semantic similarity and the word weight value, each high-frequency word can be sorted to obtain a high-frequency word sequence. The higher the semantic similarity and the higher the word weight value of the high-frequency word, the more forward it is ranked in the high-frequency word sequence.

[0097] In step S250, according to the sorting of the high-frequency words, the keyword list is determined.

[0098] In the embodiments of the present disclosure, according to the sorting of each high-frequency word, the keywords in the keyword list are determined. This screening process can be performed based on a preset threshold or a corresponding number of high-frequency words can be selected from the high-frequency word sequence according to a preset number of keywords as keywords.

[0099] In an exemplary embodiment, in the foregoing step S150, by analyzing the feedback information, the semantic similarity and word weight value corresponding to each high-frequency word can be recalculated according to the above method, and then the high-frequency words can be reordered, thereby updating the keyword list. The process of recalculating the semantic similarity and word weight value corresponding to each high-frequency word will not be repeated here.

[0100] In an exemplary embodiment, a word meaning reconstruction process can also be introduced in the process of determining the keyword list. In this word meaning reconstruction process, the selected keywords are compared with the corresponding terms in the term library, and the keywords are semantically reconstructed based on the corresponding terms in the term library. The keywords in the keyword list are replaced or supplemented by the corresponding terms in the term library to make the keywords in the finally determined keyword list better express the word meaning of the relevant keywords and ensure the accuracy of the relevant word meaning expression.

[0101] The text generation method provided by the embodiments of the present disclosure filters out high-frequency words based on word frequency for subsequent keyword evaluation, which can improve the generation efficiency of the keyword list. By evaluating the importance of high-frequency words from different dimensions through semantic similarity and word weight value, the accuracy of keyword selection can be improved, providing high-quality keyword information for the subsequent text generation process.

[0102] Figure 3 is a flowchart of a feedback information processing process shown according to an example Figure 1 As Figure 3 shown, in the embodiments of the present disclosure, the feedback information processing process of the foregoing step S150 may include the following steps.

[0103] In step S310, a feedback text for the first output text is obtained.

[0104] In the embodiments of the present disclosure, a feedback text of the user's feedback on the first output text is obtained. The feedback information of the user on the first output text is recorded in the feedback text.

[0105] In step S320, by performing semantic analysis on the feedback text, feedback semantic information is determined.

[0106] In the embodiments of the present disclosure, a language analysis model is used to perform semantic analysis on the feedback text, extract semantic content regarding the adjustment and optimization of the output text therefrom, and form feedback semantic information. Currently, there are various semantic models that can perform semantic analysis of natural language, and the present disclosure does not limit the specific model type of the semantic analysis model.

[0107] In an exemplary embodiment, this semantic analysis process generally relies on a natural language processing (NLP) model to extract and understand the core meaning of the feedback text. First, a pre-trained semantic analysis model (such as BERT, GPT, etc.) can be used to encode the feedback text to capture its contextual semantic information; then, a sentiment analysis model (such as VADER or a Transformer-based sentiment classifier) is used to determine the sentiment tendency of the feedback (such as positive, negative, or neutral); at the same time, a topic model (such as LDA) or keyword extraction technology (such as TF-IDF or TextRank) is used to identify the core topic or keywords in the feedback. In addition, an intention recognition model can be combined to analyze the specific intention of the feedback. Finally, based on the above analysis results, the feedback semantic information is determined.

[0108] In step S330, according to the feedback semantic information, the keyword list and at least one text generation factor are updated to obtain an updated keyword list and at least one text generation factor.

[0109] In the embodiments of the present disclosure, according to the feedback semantic information determined by this recognition, the previously determined keyword list and / or at least one text generation factor are further updated, so as to obtain an updated keyword list and at least one text generation factor. It should be noted that according to the feedback semantic information, part of the information in the keyword list and at least one text generation factor can be updated. In addition, according to actual needs, the corresponding relationship between the feedback semantic information and the updated content can be adjusted, and the present disclosure does not limit the mathematical relationship between the feedback semantic information and the updated content.

[0110] Figure 4 is a flowchart of a feedback information processing process shown according to an example Figure 2 As Figure 4 shown, in the embodiments of the present disclosure, the feedback information processing process of the foregoing step S150 may include the following steps.

[0111] In step S410, the feedback text for the first output text is obtained.

[0112] In the embodiments of the present disclosure, the feedback text fed back by the user for the first output text is obtained. The feedback information of the user for the first output text is recorded in the feedback text.

[0113] In step S420, the feedback text is analyzed for its type through a feedback classification model to determine the corresponding feedback type.

[0114] In the embodiments of the present disclosure, since the problems feedback by users can usually be classified into several types of problems. Therefore, the response to this feedback task can be processed based on a classification model. That is, multiple feedback types are preset according to the types of problems feedback by users. The feedback text is analyzed for its type through a pre-trained feedback classification model to determine the corresponding feedback type of the content feedback by the user, and then different feedback processing modules corresponding to different feedback types are called for differential processing.

[0115] In an exemplary embodiment, this feedback classification and analysis process generally relies on machine learning or deep learning models to classify the feedback text. First, methods based on rules or traditional machine learning (such as SVM, Naive Bayes) can be used to preliminarily classify the feedback, or pre-trained language models (such as BERT, RoBERTa, etc.) can be used for fine-grained text classification. These models can capture the context semantic information of the feedback text, thereby improving the classification accuracy. The feedback types can include content supplementation types, logical adjustment types, or language optimization types, etc. The classification model learns the features of different feedback types through training data and maps the feedback text to the most matching type during the inference stage, providing a clear guiding direction for the subsequent feedback processing module.

[0116] In an exemplary embodiment, if the feedback text involves the lack of some technical details in the document or incomplete paragraphs, it is determined as a content supplementation type according to the feedback classification model. For example, the user feedback: "The description of the 'drone battery management mechanism' in the document is too simple." If the feedback text involves disordered paragraph order or incoherent content reasoning logic, it is determined as a logical adjustment type according to the feedback classification model. For example, the user feedback: "The paragraph order in the document is chaotic." If the feedback text involves non-standard term usage or unsmooth language expression, it is determined as a language optimization type according to the feedback classification model. For example, the user feedback: "The expression about the drone battery management part is too cumbersome and not concise enough." It should be noted that the above division of feedback types is only for exemplary illustration and does not limit the protection scope of the present disclosure.

[0117] In step S430, through the feedback processing module corresponding to the feedback type, the keyword list and at least one text generation factor are updated according to the feedback text to obtain an updated keyword list and at least one text generation factor.

[0118] In the embodiments of the present disclosure, corresponding feedback processing modules are provided according to various preset feedback types to perform targeted feedback problem processing. By calling the corresponding feedback processing module, the feedback text is processed in response, and then the previously determined keyword list and / or at least one text generation factor are updated, so as to obtain an updated keyword list and at least one text generation factor. It should be noted that for different types of feedback problems, the problem processing processes are also different. Developers can further set different feedback processing modules according to the actual needs of feedback problems, and the present disclosure does not make any limitations.

[0119] In an exemplary embodiment, for the content supplement type, the content supplement processing module can guide the user to supplement the corresponding input text content or adjust the keywords in the keyword list, so as to generate an adjusted output text based on the updated keyword list to supplement the missing relevant content.

[0120] In an exemplary embodiment, for the logic adjustment type, the logic adjustment processing module can call the aforementioned proofreading model to proofread the content of the output text, reconstruct the paragraph logic, and rearrange the document order in combination with dependency syntax analysis. Dynamically adjust the Text Graph to ensure the logical coherence between paragraphs.

[0121] In an exemplary embodiment, for the language optimization type, the language optimization processing module can adjust the corresponding text generation factors, such as the temperature factor. By adjusting the corresponding text generation factors, the diversity and fluency of the optimized content are controlled to optimize the word usage and sentence patterns.

[0122] Figure 5 is a flow chart of the feedback information processing process shown by an example Figure 3 . As Figure 5 shown, in the embodiments of the present disclosure, the feedback information processing process of the aforementioned step S150 may include the following steps.

[0123] In step S510, the first output text is processed by the proofreading model to obtain text proofreading information.

[0124] In the embodiments of the present disclosure, in addition to making adjustments based on the feedback text fed back by the user, the first output text can also be processed by the proofreading model to obtain text proofreading information. The proofreading model is used to perform logical checks, semantic optimizations, and professional validations on the output text to ensure the language fluency, term accuracy, and logical consistency of the output text.

[0125] In an exemplary embodiment, the proofreading model may include one or more proofreading sub-modules, and different proofreading sub-modules are used to proofread different types of text problems. For example, the proofreading model may include any one or more of the following proofreading sub-modules.

[0126] The logical consistency sub-module is used to detect semantic conflicts or logical contradictions (such as contradictory statements, mutually exclusive concepts). The logical consistency sub-module constructs a text logic graph (TextGraph) by using dependency syntactic analysis to check the logical relationship between paragraphs. The proofreading evaluation can be carried out through the following formula:

[0127]

[0128] where S c is the logical coherence score, and E(p i , p j ) represents the logical consistency score (calculated based on semantic relationships) between paragraphs p i and p j . N represents the total number of relationships between paragraphs.

[0129] For example, the system detects a logical conflict between paragraphs A and B. Paragraph A: "The drone uses a high-efficiency lithium battery." Paragraph B: "The drone uses a solar charging solution and does not require a lithium battery." The logical consistency sub-module identifies and marks the conflict for suggesting adjustments.

[0130] The semantic optimization sub-module is used to detect the simplicity and fluency of semantic expressions. The semantic optimization sub-module streamlines the repeated or similar-meaning content in the document by calling the pre-trained deep learning model Transformer. The semantic similarity evaluation can be carried out through the following formula:

[0131]

[0132] where S(p i , p j ) represents the semantic similarity between paragraphs p i and p j . is the feature vector representation of paragraphs p i and p j .

[0133] For example, Paragraph A: "The power system consists of two motors." Paragraph B: "The power system is equipped with a dual-motor design." The semantic optimization sub-module simplifies the semantic expression: "The power system adopts a dual-motor design."

[0134] The term standardization sub-module is used to check whether the terms used are non-standard or redundant. This term standardization sub-module conducts semantic similarity comparison between the terms used in the output text and those in the term library to perform relevant term standardization review. The term standardization replacement can be carried out through the following formula:

[0135]

[0136] where T fina1 represents the finally determined term. T i represents the candidate term in the term library. S(t, T i ) represents the semantic similarity between the term t in the output text and the candidate term T i in the term library.

[0137] For example, the user inputs: "Motor energy efficiency". The term standardization sub-module replaces this term with the standard term: "Motor efficiency".

[0138] The professionalism inspection sub-module is used to inspect text errors in the output text. This professionalism inspection sub-module judges errors in the output text through multiple rule mechanisms to identify various types of errors. For example, it calls the rule engine to check common grammar errors (such as punctuation and spelling); based on the Transformer model, it identifies domain-specific inappropriate word usage or concept errors; in combination with the knowledge graph, it verifies the accuracy of technical details. The error priority can be scored through the following formula:

[0139] P e = w l ·S l + w s ·S s

[0140] where P e represents the error priority score. S1 represents the severity score of logical errors. S s represents the severity score of grammar errors. w1, w s represent the corresponding weight coefficients, which can be dynamically adjusted according to the document type.

[0141] For example, the professionalism inspection sub-module discovers and corrects the problem of inconsistent terms, and corrects "Lithium battery energy consumption" to "Lithium battery energy consumption".

[0142] It should be noted that the above-provided various review sub-modules are only for illustrative purposes and are not used to limit the protection scope of the present disclosure. Those skilled in the art can select one or several of the review sub-modules according to actual needs, or can design other dimensions of review sub-modules according to needs, and all should be regarded as within the protection scope of the present disclosure.

[0143] In addition, the proofreading module can automatically trigger the proofreading of the output text after the text generation model generates the output text, can trigger the proofreading module to proofread the output text based on the system process design, or can determine the corresponding feedback type based on the foregoing feedback classification model, and then trigger the feedback processing module to call the proofreading module to proofread the output text. The above triggering methods of the proofreading module shall all be regarded as within the protection scope of the present disclosure.

[0144] In step S520, according to the text proofreading information, update the keyword list and at least one text generation factor to obtain an updated keyword list and at least one text generation factor.

[0145] In the embodiments of the present disclosure, according to the text proofreading information, further update the determined keyword list and / or at least one text generation factor, so as to obtain an updated keyword list and at least one text generation factor. It should be noted that according to the text proofreading information, part of the information in the keyword list and at least one text generation factor can be updated. In addition, according to actual needs, the corresponding relationship between the text proofreading information and the updated content can be adjusted, and the present disclosure does not limit the mathematical relationship between the text proofreading information and the updated content.

[0146] Figure 6 is the flow of a feedback information processing process shown by an example Figure 4 . As Figure 6 shown, in the embodiments of the present disclosure, the feedback information processing process of the foregoing step S150 may include the following steps.

[0147] In step S610, obtain the feedback text for the first output text.

[0148] In the embodiments of the present disclosure, obtain the feedback text of the user's feedback on the first output text. The feedback information of the user for the first output text is recorded in the feedback text.

[0149] In step S620, by analyzing and processing the feedback text, obtain at least one feedback problem.

[0150] In the embodiments of the present disclosure, since the user may feedback multiple problems in the feedback text at the same time. Therefore, the feedback text can be analyzed and processed by a language analysis model, so as to split the problems feedback by the user to obtain multiple different feedback problems. At present, there are already a variety of semantic models that can perform semantic analysis of natural language, and the present disclosure does not limit the specific model type of the semantic analysis model.

[0151] In step S630, determine the feedback priority corresponding to each of the feedback problems.

[0152] In the embodiments of the present disclosure, for each feedback problem, its corresponding feedback priority is determined, and then processing is performed based on the priorities of different feedback problems. According to the requirements of the actual application scenario, the priorities of each feedback problem can be evaluated from multiple different dimensions, and the present disclosure does not limit the specific priority evaluation method.

[0153] In an exemplary embodiment, the feedback priorities of different feedback problems can be evaluated based on the urgency of the relevant feedback problems in the user feedback text and the frequency of feedback of the same type of problems. The feedback priority can be evaluated by the following formula:

[0154] P f = w e ·F e + w t ·F t

[0155] Wherein, P f represents the feedback priority evaluation value, and the higher the value, the higher the priority. F e represents the urgency of the user feedback. F t represents the historical occurrence frequency of the same type of feedback. w e and w t represent the weight coefficients, which are used to dynamically balance the priority.

[0156] In step S640, according to the feedback priority, the corresponding feedback problems are analyzed and processed in sequence, and the keyword list and at least one text generation factor are updated to obtain an updated keyword list and at least one text generation factor.

[0157] In the embodiments of the present disclosure, according to the feedback priorities corresponding to the above-mentioned various feedback problems, the corresponding feedback problems are analyzed and processed in sequence, and then the previously determined keyword list and / or at least one text generation factor are updated, so as to obtain an updated keyword list and at least one text generation factor. This analysis and processing process can be performed by using any of the Figures 3 to 5 above-mentioned feedback information processing methods, which will not be repeated here. By processing according to the different priorities of the feedback problems, the feedback problems with higher priorities are processed first, so as to increase the amplitude of the corresponding feedback problems in the optimization and adjustment process of the output text, thereby ensuring that the optimization process converges as early as possible and obtaining an output text that satisfies the user.

[0158] The text generation method provided by the embodiments of the present disclosure provides rich implementation manners for updating the keyword list and text generation factors based on feedback information. Through different feedback information processing methods, relevant parameters can be adjusted from different dimensions, so that the feedback iteration process is more efficient and accurate, which can greatly improve the text generation efficiency and shorten the writing time of technical documents.

[0159] Figure 7 is a flowchart of the input text supplement process shown according to an example. As Figure 7 shown, in the embodiments of the present disclosure, the process of obtaining the input text in the foregoing step S110 may further include the following steps.

[0160] In step S710, a text template that matches the input text is determined; at least one information element corresponds to the text template.

[0161] In the embodiments of the present disclosure, since there may be some necessary information elements missing in the input text provided by the user, the quality of the generated output text may be affected. Therefore, in the process of obtaining the input text, a process of guiding the user to supplement the input text by the system is further included.

[0162] In the embodiments of the present disclosure, a plurality of text templates are preset. The text template is a template of the output text preset in advance for different technical fields or different document types, and is used to standardize the format and content of the output text. At least one information element corresponds to the text template. The information element is used to represent the information type required to generate the output text corresponding to the text template. For example, document name, technical background introduction, technical problems to be solved, solution introduction, technical parameter description, etc. These information elements are usually essential for generating the output text.

[0163] In an exemplary embodiment, the corresponding text template may be matched according to the technical field involved in the input text or the document type of the output text selected by the user.

[0164] In step S720, according to the input text and at least one information element, the number of information elements covered by the input text is determined.

[0165] In the embodiments of the present disclosure, the input text can be semantically analyzed through a language analysis model, and then the text content in the input text is matched with each information element in the text template, so as to determine the number of information elements covered by the existing input text, that is, the corresponding content of the relevant information elements has been provided in the existing input text.

[0166] In an exemplary embodiment, during the process of matching the text content with each information element, a classification model can be used to perform relevant matching judgments. When it is recognized that the relevant text content belongs to the type of the information element, it is determined that the information element is covered. Conversely, when it is not recognized that the text content belongs to the type of the information element, it is determined that the information element is not covered.

[0167] In an exemplary embodiment, the keywords in the keyword list determined in the foregoing step S120 can be matched with each information element to determine whether the information element is covered. Because the subsequent text generation model mainly generates text based on the keywords in the keyword list, it is also possible to directly perform matching judgments based on the keywords in the keyword list.

[0168] In step S730, according to the number of covered information elements, an information integrity evaluation value of the input text is determined.

[0169] In the embodiments of the present disclosure, based on the number of covered information elements determined above and the number of all information elements included in the text template, the information integrity of the input text is evaluated to obtain a corresponding information integrity evaluation value. The information integrity can be evaluated by the following formula:

[0170]

[0171] where S m represents the information integrity evaluation value. The closer its value is to 1, the more serious the information loss is. n represents the number of covered information elements. h represents the number of all information elements included in the text template. The information loss degree of the input text can be evaluated through this information integrity evaluation value.

[0172] In step S740, in response to the information integrity evaluation value not reaching a preset integrity threshold, a guiding question is issued; the guiding question is used to guide the user to supplement the uncovered information elements.

[0173] In the embodiments of the present disclosure, in response to the information integrity evaluation value not reaching a preset integrity threshold, it indicates that the information loss degree in the existing input text is relatively serious. Based on this, a guiding question will be issued to the user to guide the user to supplement the uncovered information elements.

[0174] In an exemplary embodiment, according to the information elements missing by the user, the text content of corresponding guiding questions is generated by a guiding question generation model, so as to better guide the user to provide relevant content. Here, the guiding question generation model is essentially a text generation model. It can directly use the text generation model used in the foregoing step S140 to generate guiding questions, or a dedicated guiding question generation model can be separately trained based on a large language model to generate guiding questions. Similar to the foregoing S140, the process of the guiding question generation model generating guiding questions can also be parameter-adjusted through relevant adjustment factors. Specifically, it may include the following steps.

[0175] In response to the information integrity evaluation value not reaching a preset integrity threshold, determine at least one guiding question generation factor according to the information integrity evaluation value;

[0176] Generate and send the guiding question through the guiding question generation model according to the uncovered information elements and at least one guiding question generation factor.

[0177] In the embodiments of the present disclosure, the guiding question generation factor can be determined according to the information integrity evaluation value S determined above. m This guiding question generation factor can refer to the foregoing text generation factors, including one or more of a temperature factor, a repetition penalty factor, and a semantic adaptation factor, which will not be repeated here. It should be noted that this guiding question generation factor and the foregoing text generation factor are independent adjustment factors, and are only used for parameter adjustment of generating guiding questions.

[0178] Similar to the text generation process of the foregoing text generation model, the guiding question is generated and sent through the guiding question generation model according to the determined uncovered information elements and the above guiding question generation factors. The generation of the guiding question can be controlled by the following formula:

[0179]

[0180] Among them, P(q): the generation probability of the guiding question q. E(q): the semantic relevance score of the question, calculated by the generation model.

[0181] In an exemplary embodiment, the user inputs the text: "automatic driving system". The system detects that there are missing information elements, for example, missing information such as a perception module and a planning algorithm. The information integrity evaluation value S is calculated through the above formula. m= 0.67, which does not reach the preset integrity threshold, triggering the generation of guiding questions. The system generates guiding questions: "Does the autonomous driving system include a perception module? What is the main function of the perception module?" The user supplements the input text: "Yes, it includes a lidar and a camera perception module." Based on the user's supplementary content, the system recalculates the information integrity evaluation value S m = 0.33. If this information integrity evaluation value reaches the preset integrity threshold, the generation of guiding questions stops. If this information integrity evaluation value still does not reach the preset integrity threshold, the generation of guiding questions continues.

[0182] In an exemplary embodiment, during the process of generating guiding questions, if multiple guiding questions need to be generated, the priority of the guiding questions can be further evaluated. The system generates corresponding guiding questions in sequence according to this priority. The priority of the guiding questions can be evaluated through the following formula, combining the information integrity evaluation value S m and the importance weight value W of information elements k to evaluate the priority of the guiding questions:

[0183] P g = S m ·W k

[0184] where W k : the importance weight value of the currently missing information element in the technical field, which can be provided by the term library.

[0185] The text generation method provided by the embodiments of the present disclosure can, by judging the information integrity of the input text, actively generate guiding questions when the information input by the user is insufficient, helping the user to supplement key information, thereby improving the integrity and accuracy of document generation.

[0186] The following are the device embodiments of the present disclosure, which can be used to execute the method embodiments of the present disclosure. For the details not disclosed in the device embodiments of the present disclosure, please refer to the method embodiments of the present disclosure.

[0187] Figure 8 is a block diagram of a text generation device shown according to an exemplary embodiment. Referring to Figure 8 , the device 800 may include: an input text acquisition unit 810, a keyword determination unit 820, a text generation factor determination unit 830, a first text generation unit 840, a feedback information unit 850, and a second text generation unit 860.

[0188] The input text acquisition unit 810 is configured to acquire the input text input by the user and the term library corresponding to the input text.

[0189] A keyword determination unit 820, configured to determine a keyword list by performing semantic analysis on the input text; the keyword list includes at least one keyword.

[0190] A text generation factor determination unit 830, configured to determine at least one text generation factor; the text generation factor is used to adjust text generation parameters.

[0191] A first text generation unit 840, configured to generate a first output text according to the term library, the keyword list, and at least one text generation factor through a pre-trained text generation model.

[0192] A feedback information unit 850, configured to obtain feedback information on the first output text; by analyzing and processing the feedback information, update the keyword list and at least one text generation factor to obtain an updated keyword list and at least one text generation factor.

[0193] A second text generation unit 860, configured to generate a second output text according to the term library, the updated keyword list, and at least one text generation factor through the text generation model.

[0194] In some exemplary embodiments of the present disclosure, the keyword determination unit 820 is further configured to perform word frequency analysis on the input text, screen out at least one high-frequency word; calculate the semantic similarity between each high-frequency word and the input text through a semantic analysis model; determine the word weight value corresponding to each high-frequency word; sort the high-frequency words in the at least one high-frequency word according to the semantic similarity and the word weight value; determine the keyword list according to the sorting of the high-frequency words.

[0195] In some exemplary embodiments of the present disclosure, the at least one text generation factor includes at least one of the following text generation factors: a temperature factor, used to adjust the randomness of the generated content; a repetition penalty factor, used to adjust the generation probability of repeated content; a semantic adaptation factor, used to adjust the matching degree between the generated content and the term library.

[0196] In some exemplary embodiments of the present disclosure, the feedback information unit 850 is further configured to obtain a feedback text on the first output text; determine feedback semantic information by performing semantic analysis on the feedback text; update the keyword list and at least one text generation factor according to the feedback semantic information to obtain an updated keyword list and at least one text generation factor.

[0197] In some exemplary embodiments of the present disclosure, the feedback information unit 850 is further configured to obtain the feedback text for the first output text; perform type analysis on the feedback text through a feedback classification model to determine the corresponding feedback type; and update the keyword list and at least one text generation factor according to the feedback text through a feedback processing module corresponding to the feedback type, so as to obtain an updated keyword list and at least one text generation factor.

[0198] In some exemplary embodiments of the present disclosure, the feedback information unit 850 is further configured to perform text proofreading processing on the first output text through a proofreading model to obtain text proofreading information; and update the keyword list and at least one text generation factor according to the text proofreading information, so as to obtain an updated keyword list and at least one text generation factor.

[0199] In some exemplary embodiments of the present disclosure, the feedback information unit 850 is further configured to obtain the feedback text for the first output text; perform analysis processing on the feedback text to obtain at least one feedback problem; determine the feedback priority corresponding to each feedback problem; and update the keyword list and at least one text generation factor by sequentially performing analysis processing on the corresponding feedback problems according to the feedback priority, so as to obtain an updated keyword list and at least one text generation factor.

[0200] In some exemplary embodiments of the present disclosure, the input text acquisition unit 810 is further configured to determine a matching text template according to the input text; at least one information element corresponds to the text template; determine the number of information elements covered by the input text according to the input text and at least one information element; determine an information integrity evaluation value of the input text according to the number of covered information elements; and issue a guiding question in response to the information integrity evaluation value not reaching a preset integrity threshold; the guiding question is used to guide the user to supplement the uncovered information elements.

[0201] In some exemplary embodiments of the present disclosure, the input text acquisition unit 810 is further configured to determine at least one guiding question generation factor according to the information integrity evaluation value in response to the information integrity evaluation value not reaching a preset integrity threshold; and generate and issue the guiding question through a guiding question generation model according to the uncovered information elements and at least one guiding question generation factor.

[0202] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0203] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory including instructions, and the above instructions can be executed by a processor of the device to complete the above method. Optionally, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0204] In an exemplary embodiment, a computer program product is also provided, including a computer program which, when executed by a processor, implements the method in the above embodiment.

[0205] After considering the specification and practicing the disclosed invention herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed herein. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0206] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A text generation method, characterized in that, including: obtaining input text entered by a user and a term library corresponding to the input text; determining a keyword list by performing semantic analysis on the input text; the keyword list includes at least one keyword; determining at least one text generation factor; the text generation factor is used to adjust text generation parameters; generating a first output text by a pre-trained text generation model according to the term library, the keyword list, and at least one text generation factor; obtaining feedback information on the first output text; by analyzing and processing the feedback information, updating the keyword list and at least one text generation factor to obtain an updated keyword list and at least one text generation factor; generating a second output text by the text generation model according to the term library, the updated keyword list, and at least one text generation factor.

2. The method according to claim 1, wherein The determining a keyword list by performing semantic analysis on the input text includes: performing word frequency analysis on the input text to screen out at least one high-frequency word; calculating the semantic similarity between each high-frequency word and the input text through a semantic analysis model; determining the word weight value corresponding to each high-frequency word; sorting the high-frequency words in the at least one high-frequency word according to the semantic similarity and the word weight value; determining the keyword list according to the sorting of the high-frequency words.

3. The method according to claim 1, characterized in that The at least one text generation factor includes at least one of the following text generation factors: a temperature factor, which is used to adjust the randomness of the generated content; a repetition penalty factor, which is used to adjust the generation probability of repeated content; a semantic adaptation factor, which is used to adjust the matching degree between the generated content and the term library.

4. The method according to claim 1, wherein The obtaining feedback information on the first output text; by analyzing and processing the feedback information, updating the keyword list and at least one text generation factor to obtain an updated keyword list and at least one text generation factor includes: obtaining feedback text on the first output text; determining feedback semantic information by performing semantic analysis on the feedback text; updating the keyword list and at least one text generation factor according to the feedback semantic information to obtain an updated keyword list and at least one text generation factor.

5. The method according to claim 1, characterized in that, The obtaining feedback information on the first output text; by analyzing and processing the feedback information, updating the keyword list and at least one text generation factor to obtain an updated keyword list and at least one text generation factor includes: obtaining the feedback text on the first output text; performing type analysis on the feedback text through a feedback classification model to determine the corresponding feedback type; updating the keyword list and at least one text generation factor according to the feedback text through a feedback processing module corresponding to the feedback type to obtain an updated keyword list and at least one text generation factor.

6. The method according to claim 1, characterized in that, Obtaining feedback information on the first output text; updating the keyword list and at least one text generation factor by analyzing and processing the feedback information to obtain an updated keyword list and at least one text generation factor, including: Performing text proofreading processing on the first output text through a proofreading model to obtain text proofreading information; Updating the keyword list and at least one text generation factor according to the text proofreading information to obtain an updated keyword list and at least one text generation factor.

7. The method according to claim 1, wherein Obtaining feedback information on the first output text; updating the keyword list and at least one text generation factor by analyzing and processing the feedback information to obtain an updated keyword list and at least one text generation factor, including: Obtaining the feedback text on the first output text; Obtaining at least one feedback problem by analyzing and processing the feedback text; Determining the feedback priority corresponding to each feedback problem; Updating the keyword list and at least one text generation factor by sequentially analyzing and processing the corresponding feedback problems according to the feedback priority to obtain an updated keyword list and at least one text generation factor.

8. The method according to claim 1, characterized in that, The obtaining of the input text input by the user and the term library corresponding to the input text further includes: Determining a matching text template according to the input text; the text template corresponds to at least one information element; Determining the number of information elements covered by the input text according to the input text and at least one information element; Determining an information integrity evaluation value of the input text according to the number of covered information elements; In response to the information integrity evaluation value not reaching a preset integrity threshold, issuing a guiding question; the guiding question is used to guide the user to supplement the uncovered information elements.

9. The method according to claim 8, wherein The issuing of the guiding question in response to the information integrity evaluation value not reaching the preset integrity threshold includes: In response to the information integrity evaluation value not reaching the preset integrity threshold, determining at least one guiding question generation factor according to the information integrity evaluation value; Generating and issuing the guiding question through a guiding question generation model according to the uncovered information elements and at least one guiding question generation factor.

10. A text generation device, characterized in that, Including: An input text acquisition unit, configured to acquire the input text input by the user and the term library corresponding to the input text; A keyword determination unit, configured to determine a keyword list by performing semantic analysis on the input text; the keyword list includes at least one keyword; A text generation factor determination unit, configured to determine at least one text generation factor; The text generation factor is used to adjust text generation parameters; A first text generation unit, configured to generate a first output text through a pre-trained text generation model according to the term library, the keyword list, and at least one text generation factor; A feedback information unit, configured to obtain feedback information on the first output text; by analyzing and processing the feedback information, update the keyword list and at least one text generation factor to obtain an updated keyword list and at least one text generation factor; A second text generation unit, configured to generate a second output text according to the term library, the updated keyword list and at least one text generation factor through the text generation model.