Tibetan news generation method, system and equipment and storage medium
By using TF-IDF to filter prompt words in Tibetan news generation, combining convolutional neural network and Transformer model, and using dynamic weighted fusion and hybrid decoding strategies, the problems of lack of details and logic in Tibetan news generation are solved, and higher quality Tibetan news generation is achieved.
Patent Information
- Application Number
- CN202510129449.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-05-30
AI Technical Summary
The existing technology has problems such as lack of details, inaccuracy, improper logic or duplication and redundancy in Tibetan news generation, especially in dealing with Tibetan characteristics and in-depth language research.
By obtaining Tibetan news corpus data, extracting prompt words in the title, calculating their TF-IDF weights and filtering, adding convolutional neural networks and Transformer models for training, combining dynamic weighted fusion modules and hybrid decoding strategies, an improved Transformer model is generated to improve the quality of news generation.
The generated Tibetan news content is more accurate, logical and fluent in language, solving the problems of theme deviation and details in the generation of long texts by the traditional Transformer model, and reducing the generation of redundant and meaningless content.
Smart Images

Figure CN120068811A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information processing, and particularly relates to a Tibetan news generation method, system, device, and storage medium. Background Art
[0002] Due to its unique grammatical structure, complex language rules, and limited corpus size, Tibetan still faces many challenges in current Tibetan news generation. Compared with languages with rich corpus resources, the generated Tibetan content often lacks details, is inaccurate, and even has illogical or redundant phenomena. This limitation not only stems from the deficiencies of the model in dealing with Tibetan characteristics but also reflects the lack of in-depth research on the Tibetan language.
[0003] Since the advent of Transformer, generative text technology has been widely used in various fields. Especially in the field of automated news writing, its efficient information processing ability has significantly improved the efficiency of news generation. This technology can not only quickly generate news content with clear logic and fluent language but also effectively alleviate the limitations of traditional news production methods in terms of manpower and time. However, it exposes obvious deficiencies when dealing with Tibetan news generation. First, the Transformer encoder has limited ability to capture fine-grained local features of Tibetan texts, resulting in generated texts lacking details and semantic accuracy. Second, the content generated by its decoder often appears repetitive or unnatural, affecting the fluency and overall quality of the text. These problems restrict the practical application of automated Tibetan news generation, especially in scenarios that require high-quality output. Based on this, how to develop an innovative model for the special needs of Tibetan news generation has become an urgent problem to be solved. Summary of the Invention
[0004] To solve the deficiencies in the existing methods for generating Tibetan news, the present invention provides a Tibetan news generation method, system, device, and storage medium.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A Tibetan news generation method, comprising the following steps:
[0007] Obtain Tibetan news corpus data, and extract prompting words according to the title data in the Tibetan news corpus data;
[0008] Calculate the TF-IDF weights of the prompting words in the Tibetan news corpus data, screen the prompting words according to a set weight threshold, and splice the screened prompting words to the corresponding positions of the Tibetan news corpus data;
[0009] A convolutional neural network is added after the embedding layer of the original Transformer model. A dynamic weighted fusion module is introduced after the convolutional neural network and the Transformer encoder, and a hybrid decoding strategy is adopted in the decoder to form an improved Transformer model. The concatenated news corpus data is input into the improved Transformer model to train the model, and a news generation model is obtained.
[0010] The text to be detected is input into the news generation model. The convolutional neural network and the Transformer encoder respectively obtain the local features and global semantic information of the text to be detected. The local features and global semantic information are feature fused through the dynamic weighted fusion module to obtain a feature sequence. The feature sequence is input into the decoder and decoded through the hybrid decoding strategy to output the Tibetan news content.
[0011] Preferably, the convolutional neural network performs local feature extraction on the prompt word through one-dimensional convolution operation, specifically:
[0012] C local = GeLU(Conv1D(E' pw , k, s, p));
[0013] Where Conv1D is a one-dimensional convolution function; GeLU is an activation function; k represents the convolution kernel; p represents padding; s represents the stride; E' pw represents the prompt word vector.
[0014] Preferably, the feature fusion of the local features and global semantic information through the dynamic weighted fusion module is specifically through the following formula:
[0015] C fusion = α·T global + β·C local ;
[0016] Where α and β are dynamically learned weight coefficients; T global is the global semantic feature; C local is the local feature.
[0017] Preferably, the hybrid decoding strategy specifically includes: successively adopting the Top-k strategy, the Top-p strategy, temperature adjustment, and polynomial sampling for decoding.
[0018] Preferably, calculate the TF-IDF weight of the prompt word in the Tibetan news corpus data, screen the prompt word according to the set weight threshold, and splice the screened prompt word to the corresponding position of the Tibetan news corpus data, specifically including the following steps:
[0019] Calculate the term frequency TF of the prompt word in the target document;
[0020] Calculate the inverse document frequency IDF of the prompt word, specifically through the following formula:
[0021]
[0022] Where M is the total number of documents in the corpus and N is the number of documents containing the prompt word; the TF-IDF value of the prompt word is the product of TF and IDF;
[0023] Filter out the prompt words to be spliced according to the preset TF-IDF threshold;
[0024] Sort the filtered prompt words from largest to smallest and splice them into the corresponding news corpus in turn to form a complete spliced corpus.
[0025] Preferably, the processing process of the embedding layer in the news generation model specifically includes the following steps:
[0026] Convert the spliced corpus into low-dimensional word vectors through word embedding;
[0027] Extract the weight feature Weigt(w i ) and position feature Position(w i ) for each word in the low-dimensional word vector;
[0028] Encode through the following formula to obtain a token sequence with semantic information;
[0029] Token i =f(Embedding(w i ),Weigt(w i ),Position(w i ));
[0030] Input the token sequence into a convolutional neural network and a Transformer encoder.
[0031] The present invention also provides a Tibetan news generation system based on a title and a prompt word, specifically including:
[0032] A data acquisition module for acquiring Tibetan news corpus data and extracting prompt words according to the title data in the Tibetan news corpus data.
[0033] A corpus processing module for calculating the TF-IDF weights of the prompt words in the Tibetan news corpus data, filtering the prompt words according to the set weight threshold, and splicing the filtered prompt words to the corresponding positions of the Tibetan news corpus data.
[0034] A model construction module is used to add a convolutional neural network after the embedding layer of the original Transformer model. A dynamic weighted fusion module is introduced after the convolutional neural network and the Transformer encoder, and a hybrid decoding strategy is adopted in the decoder to construct an improved Transformer model. The concatenated news corpus data is input into the improved Transformer model to train the model, and a news generation model is obtained.
[0035] A news generation module is used to input the text to be detected into the news generation model. The convolutional neural network and the Transformer encoder respectively obtain the local features and global semantic information of the text to be detected, and perform feature fusion on the local features and global semantic information through the dynamic weighted fusion module to obtain a feature sequence. The feature sequence is input into the decoder and decoded through the hybrid decoding strategy to output the Tibetan news content.
[0036] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps in the method for generating Tibetan news.
[0037] The present invention also provides a computer-readable storage medium. A computer program is stored on the storage medium. When the computer program is loaded by the processor, it can execute the steps in the method for generating Tibetan news.
[0038] The method for generating Tibetan news provided by the present invention has the following beneficial effects:
[0039] In the present invention, relevant prompt words are extracted from the news titles of the Tibetan news corpus, and the weights of the prompt words are calculated to screen the prompt words, guiding the model to generate more accurate, logical, and fluent Tibetan news texts. Adding a convolutional neural network to the original Transformer model further strengthens the ability to capture local features of the Tibetan news text, helps the model extract more detailed information, and introduces a dynamic weighted fusion module for feature fusion after the convolutional neural network and the Transformer encoder are processed, solving the problems of theme deviation and detail loss existing in the traditional Transformer model in long text generation. Introducing a hybrid decoding strategy in the decoding process of the Transformer model effectively reduces the generation of redundant words and meaningless content, avoiding the problems of repetition and logical incoherence in the traditional Transformer model. The generated text is more natural and fluent, meeting the requirements of news dissemination. Description of the Drawings
[0040] To more clearly illustrate the embodiments of the present invention and their design solutions, the accompanying drawings required for this embodiment will be briefly introduced below. The accompanying drawings in the following description are only partial embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0041] Figure 1 It is a flowchart of a Tibetan news generation method of the present invention.
[0042] Figure 2 It is a structural diagram of an improved Transformer model of the present invention.
[0043] Figure 3 It is a schematic flowchart of the decoder of the Transformer model in the embodiment of the present invention.
[0044] Figure 4 It is a schematic flowchart of the hybrid decoding strategy of the decoder in the embodiment of the present invention. Detailed implementation manners
[0045] In order to enable those skilled in the art to better understand the technical solutions of the present invention and implement them, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.
[0046] Embodiment
[0047] The present invention provides a Tibetan news generation method, as Figure 1 shown, which specifically includes the following steps:
[0048] S1. Use web crawler technology to obtain Tibetan news corpus data of different fields and types from databases such as Tibetan news websites, media platforms, and news portals, including fields such as politics, economy, and culture, to ensure the diversity of the corpus. During the collection process, the title data can effectively summarize the news content and provide key information for subsequent generation tasks. According to the title data, relevant prompt words are extracted from the news corpus. The prompt words (Prompt words, pw) are important words of the news theme and can accurately reflect the core content of the news.
[0049] S2. Calculate the TF-IDF weights of the prompt words in the Tibetan news corpus data, screen the prompt words according to the set weight threshold, and splice the screened prompt words to the corresponding positions of the Tibetan news corpus data. Specifically:
[0050] Calculate the term frequency TF of the prompt word in the target document;
[0051] Calculate the inverse document frequency (IDF) of the prompt word, specifically through the following formula:
[0052]
[0053] Among them, M is the total number of documents in the corpus, and N is the number of documents containing the prompt word; the TF-IDF value of the prompt word is the product of TF and IDF;
[0054] The prompt words to be spliced are screened according to a preset TF-IDF threshold;
[0055] The screened prompt words are sorted from largest to smallest and spliced into the corresponding news corpus in turn to form the final spliced corpus.
[0056] S3. Introduce a convolutional neural network after the embedding layer of the original Transformer model, introduce a dynamic weighted fusion module after the convolutional neural network and the Transformer encoder to form a new encoder, and adopt a hybrid decoding strategy in the decoder to obtain an improved Transformer model. Input the spliced news corpus data into the improved Transformer model to train the model to obtain a news generation model, as Figure 2 shown.
[0057] Specifically, in the new encoder, the embedding layer (Embedding) is divided into word embedding and positional embedding. In the Tibetan news generation task, the title and the prompt word are used as inputs, and the input sequence X = [x 1 , x 2 , …, x n is converted into a digital sequence X id , and the corresponding word embedding is obtained by indexing the rows of the Embedding Matrix through X id . The Embedding Matrix is a matrix with a shape of V×d model , where V is the size of the vocabulary and d model is the hidden layer dimension.
[0058] The parameters in the Embedding Matrix are randomly initialized. It is convenient to adjust them through optimization algorithms such as gradient descent during the training process, so that the model can learn effective representations. The position encoding matrix (P) adds the corresponding position encoding vector p i to each position x i to capture the absolute position information of the input sequence X. The syllable embedding vector w i is added to the position encoding vector p i to obtain the input sequence E i . The formula is as follows:
[0059] W(x i ) = Embedding Matrix[x i ;
[0060]
[0061] E i = w i + p i ;
[0062] Among them, i represents the position and j represents the index of the dimension.
[0063] The improved Transformer model processes the output data of the embedding layer, specifically:
[0064] The prompt input sequence passes through a convolutional neural network (local convolution module L conv ), extracts the local semantic information in the prompt sequence, and enhances the controllability and detail expression ability of the generated content. Apply a convolutional kernel (filter) to perform a sliding window operation on the input data to extract local features. Then normalize the convolutional features to stabilize the training process and accelerate convergence. Next, apply the GeLU activation function to introduce a non-linear transformation to enhance the expression ability of the model. Finally, the pooling layer performs downsampling on the convolutional features, and reduces the size of the feature map through max pooling to obtain the local features of the input sequence. Calculate through the following formula:
[0065] C local = GeLU(Conv1D(E′ pw , k, s, p));
[0066] Among them, Conv1D is a one-dimensional convolution function; GeLU is an activation function; k represents the convolutional kernel; p represents padding; s represents the stride; E' pw represents the prompt vector.
[0067] Through the convolution operation, it can efficiently extract the phrase structure and local semantic features in the prompt, capture the close connections between prompts, especially the semantic dependency relationships between phrases. These local features play a key role in the news text generation process, and can automatically learn and complete background information, including the relationships between time, place, people and events, to ensure that the generated text is highly consistent and coherent in content.
[0068] At the same time, the title input sequence passes through the Transformer encoder (global attention module G Trm) In the news generation task, the title is often a highly condensed summary of the news content, containing key information and the core theme. To effectively model the global information of the title, the input sequence E is mapped to Query (Q), Key (K), and Value (V) through a linear transformation. The multi-head self-attention mechanism (Multi-Head Attention, MHA) is used to calculate the similarity between each word in the title and other words, and extract global semantic features. The specific implementation is as follows: i Map to Query (Q), Key (K), and Value (V). Use the multi-head self-attention mechanism (Multi-Head Attention, MHA) to calculate the similarity between each word in the title and other words, and extract global semantic features. The specific implementation is as follows:
[0069] Q = E' title ·W Q , K = E' title ·W K , V = E' title ·W V ;
[0070]
[0071] Among them, is the scaling factor of the attention score; E' title represents the title vector; W Q , W K and W V represent the weight matrices respectively; b global is the bias of the global attention module.
[0072] The global features of the title can constrain the generated text, effectively reducing content redundancy, information deviation, and logical errors caused by long text generation, and ensuring that the generated text is highly consistent with the title semantically. After each Multi-Head Attention layer, there is a feed-forward fully connected layer (Feed-Forward Neural Network, FFN). The feed-forward neural network realizes the non-linear mapping of features through the combination of two linear transformations and the GeLU activation function. The specific calculation formula is as follows:
[0073] FNN(E' title ) = GeLU(W 2 (ReLU(W 1 E' title +b 1 ))+b 2 ).
[0074] After being processed by MHA and the previous FFN, the model can accurately identify the keywords, core phrases, and overall theme structure in the title, strengthen the expression effect of the title features, and provide a clear theme framework and global semantic guidance for the news generation task.
[0075] The extracted global features and local features are weighted and fused through a dynamic weighted fusion module to form a unified context representation. The fused features are represented as:
[0076] C fusion = α · T global + β · C local ;
[0077] Among them, α is a dynamically learned weight coefficient. Through the dynamic weighted fusion mechanism, the model can automatically adjust the weights of global features and local features according to the requirements of the generation task, ensuring that the generated text is more realistic and credible in terms of details while maintaining the coherence of the overall theme. In addition, it can also control the degree of dependence on the prompt words during the generation process, focus on the expression of key information, and avoid generating irrelevant or redundant information.
[0078] The decoder uses the key (K) and value (V) matrices output by the new encoder to calculate the attention weights between the query (Q) matrix and the key matrix. The output of the new encoder includes the title and prompt words, which provide the global information and important prompts of the input text as the feature representations of the key and value.
[0079] The decoder can focus on the output of the new encoder related to the current generation position, thus ensuring the consistency and relevance of the generated text with the content of the source text. For example, when processing a title " (The Global Climate Change Conference was held in Paris)" and related prompt words such as " (Low-carbon cycle)", " (Ecological civilization construction)", " (International cooperation)", when the decoder generates news article content related to the conference, it will use Encoder-Decoder Attention to focus on the importance and relevance of these keywords in the whole article. This ensures that the generated text content closely revolves around the conference theme and avoids the appearance of contradictions and irrelevant content during the generation process.
[0080] A hybrid decoding strategy is adopted in the decoder, such as Figure 4As shown, first, through the Top-k strategy, the top k words with the highest probabilities (k = 3) are retained in the probability distribution of the generated vocabulary. By restricting the number of candidate words, the Top-k strategy reduces the occurrence of low-probability words and improves the coherence and relevance of the generated text. Next, the Top-p strategy is used to retain the words whose cumulative probabilities in the probability distribution of the generated vocabulary are higher than the threshold (p = 0.8). The Top-p strategy dynamically adjusts the number of candidate words, ensuring that more important words are covered during the generation process while avoiding the generation of rare words, thereby increasing the diversity and creativity of the generated text. Then, the model performs temperature adjustment on the probability distribution filtered by Top-k and Top-p. By adjusting the temperature parameter (Temperature = 0.7), the probability distribution of the generated text is normalized. A lower temperature value makes the model tend to select high-probability words, improving the quality and consistency of the generated text, endowing the model with moderate creativity, and making the generated text richer and more diverse. Finally, the model performs multinomial sampling according to the adjusted probability distribution to obtain the index of the next word. Multinomial sampling introduces a certain degree of randomness, avoiding the monotonicity and repetition of the generated text and making the final text more natural and in line with human expression habits.
[0081] The hybrid decoding strategy shows significant advantages in the Tibetan news generation task. It not only improves the quality and fluency of the generated text but also ensures the coherence and relevance of the text. By moderately introducing randomness and diversity, this strategy effectively enhances the creativity and naturalness of the generated text, making it more in line with the language characteristics of actual news reports.
[0082] S4. Input the text to be detected into the news generation model and output the Tibetan news content.
[0083] The beneficial effects of the present invention include: (1) By combining the title and prompt words, this method guides the model to generate more accurate, logical, and fluent Tibetan news texts. Compared with traditional text generation methods, it can improve the detail richness and accuracy of the generated content on limited Tibetan corpus data, significantly reducing the repetition and unnatural phenomena that occur during the generation process. (2) By introducing an architecture that combines a convolutional neural network and a Transformer, the ability to capture local features of Tibetan news texts is further enhanced. The convolutional layer helps the model extract more detailed information, resulting in more delicate and realistic news content, improving the accuracy of the text and the richness of expression. (3) During the decoding process, the hybrid decoding strategy effectively reduces the generation of redundant words and meaningless content, avoiding the repetition and logical incoherence problems commonly found in traditional Transformer models. The generated text is more natural and fluent, meeting the requirements of news dissemination. (4) The present invention provides a new direction for technological innovation in the field of Tibetan news generation, overcomes the bottlenecks in Tibetan generation of existing technologies, promotes the application of automated news writing technology in the Tibetan media industry, and provides new possibilities for future Tibetan news production and dissemination.
[0084] The present invention also provides a Tibetan news generation system, specifically including:
[0085] A data acquisition module, used to acquire Tibetan news corpus data and extract prompt words according to the title data in the Tibetan news corpus data.
[0086] A corpus processing module, used to calculate the TF-IDF weights of the prompt words in the Tibetan news corpus data, screen the prompt words according to the set weight threshold, and splice the screened prompt words to the corresponding positions of the Tibetan news corpus data.
[0087] A model construction module, used to add a convolutional neural network after the embedding layer of the original Transformer model, introduce a dynamic weighted fusion module after the convolutional neural network and the Transformer encoder, and adopt a hybrid decoding strategy in the decoder to form an improved Transformer model; input the spliced news corpus data into the improved Transformer model to train the model and obtain a news generation model.
[0088] A news generation module, used to input the text to be detected into the news generation model. The convolutional neural network and the Transformer encoder respectively obtain the local features and global semantic information of the text to be detected, perform feature fusion on the local features and global semantic information through the dynamic weighted fusion module to obtain a feature sequence; input the feature sequence into the decoder and perform decoding through the hybrid decoding strategy to output Tibetan news content.
[0089] Each module in the above-mentioned Tibetan news generation system based on titles and prompt words can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in the processor of a computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0090] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps in an embodiment of a Tibetan news generation method. For the specific implementation method, reference can be made to the method embodiment, which will not be elaborated here.
[0091] Furthermore, the present invention also provides a non-transitory computer-readable storage medium containing instructions, on which a computer program is stored. For example, a memory containing instructions. The above instructions can be executed by the processor of a computer device to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc. When the computer program is executed by the processor, it can implement the steps in an embodiment of a Tibetan news generation method. For the specific implementation method, reference can be made to the method embodiment, which will not be elaborated here.
[0092] Those skilled in the art should understand that the embodiments of the present invention can provide methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0093] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can also be implemented. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0094] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more processes and / or blocks Figure 1 in the process Figure 1 or processes and / or boxes
[0095] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more processes and / or blocks Figure 1 in the process Figure 1 or processes and / or boxes
[0096] It should be noted that the above-described specific embodiments may enable those skilled in the art to more fully understand the present invention, but do not limit the present invention in any way. Therefore, although the present specification and embodiments have described the present invention in detail, those skilled in the art should understand that the present invention may still be modified or equivalently replaced; and all technical solutions and improvements that do not depart from the spirit and scope of the present invention are covered by the protection scope of the patent of the present invention. Any reference numeral in a claim should not be construed as limiting the claim involved. Any simple variation or equivalent replacement of a technical solution that can be obviously obtained by any person skilled in the art within the technical scope disclosed by the present invention falls within the protection scope of the present invention.
Claims
1. A Tibetan news generation method, characterized in that: The following steps are involved: Acquire Tibetan news corpus data, and extract prompt words according to title data in the Tibetan news corpus data; Calculating the TF-IDF weight of the prompt word in the Tibetan news corpus data, screening the prompt word according to a set weight threshold, and splicing the screened prompt word to a corresponding position of the Tibetan news corpus data; A convolutional neural network is added after the embedding layer of the original Transformer model, a dynamic weighted fusion module is introduced after the convolutional neural network and the Transformer encoder, and a hybrid decoding strategy is adopted in the decoder to form an improved Transformer model; the concatenated news corpus data is input into the improved Transformer model to train the model and obtain a news generation model; The text to be detected is input into the news generation model, the local features and global semantic information of the text to be detected are obtained respectively through a convolutional neural network and a Transformer encoder, and the local features and the global semantic information are fused through a dynamic weighted fusion module to obtain a feature sequence; the feature sequence is input into a decoder, decoded through a hybrid decoding strategy, and Tibetan news content is output.
2. The Tibetan news generation method according to claim 1, characterized in that: The convolutional neural network extracts local features of the prompt word through a one-dimensional convolution operation, specifically: C local =GeLU(Conv1D(E' pw ,k,s,p)): Among them, Conv1D is a one-dimensional convolution function; GeLU is an activation function; k represents the convolution kernel; p represents padding; s represents the step size; E' pw Represents the prompt word vector.
3. The Tibetan news generation method according to claim 1, characterized in that: The local features and global semantic information are fused by the dynamic weighted fusion module, specifically by the following formula: C fusion =α·T global +β·C local ; Among them, α and β are the weight coefficients of dynamic learning; T global is the global semantic feature; C local It is a local feature.
4. The Tibetan news generation method according to claim 1, characterized in that: The hybrid decoding strategy specifically includes: sequentially adopting the Top-k strategy, the Top-p strategy, temperature adjustment and polynomial sampling for decoding.
5. The Tibetan news generation method according to claim 1, characterized in that: Calculating the TF-IDF weight of the prompt word in the Tibetan news corpus data, screening the prompt word according to a set weight threshold, and splicing the screened prompt word to the corresponding position of the Tibetan news corpus data, specifically includes the following steps: Calculate the word frequency TF of the prompt word in the target document; Calculate the inverse document frequency (IDF) of the prompt word using the following formula: Where M is the total number of documents in the corpus, and N is the number of documents containing the prompt word. The TF-IDF value of the prompt word is the product of TF and IDF. Filter out the prompt words that need to be spliced according to the preset TF-IDF threshold; The screened prompt words are sorted from largest to smallest, and are sequentially spliced into the corresponding news corpus to form a complete spliced corpus.
6. The Tibetan news generation method according to claim 1, characterized in that: The processing process of the embedding layer in the news generation model specifically includes the following steps: Convert the concatenated corpus into low-dimensional word vectors through word embedding; Extract the weight feature Weigt(w i ) and position feature Position(w i ); Encode through the following formula to obtain a token sequence with semantic information; Token i =f(Embedding(w i ),Weigt(w i ),Position(w i )); The word-unit token sequence is input into the convolutional neural network and the Transformer encoder.
7. A Tibetan news generation system, characterized in that: include: A data acquisition module, used to acquire Tibetan news corpus data, and extract prompt words according to the title data in the Tibetan news corpus data; A corpus processing module, used for calculating the TF-IDF weight of the prompt word in the Tibetan news corpus data, screening the prompt word according to a set weight threshold, and splicing the screened prompt word to the corresponding position of the Tibetan news corpus data; A model building module is used to add a convolutional neural network after the embedding layer of the original Transformer model, introduce a dynamic weighted fusion module after the convolutional neural network and the Transformer encoder, and adopt a hybrid decoding strategy in the decoder to form an improved Transformer model; input the spliced news corpus data into the improved Transformer model to train the model and obtain a news generation model; The news generation module is used to input the text to be detected into the news generation model, and the convolutional neural network and the Transformer encoder respectively obtain the local features and global semantic information of the text to be detected, and the local features and the global semantic information are fused by the dynamic weighted fusion module to obtain a feature sequence; the feature sequence is input into the decoder, and the decoded by the hybrid decoding strategy to output the Tibetan news content.
8. A computer device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is loaded into a processor, it can execute the steps of the method according to any one of claims 1 to 6.
Citation Information
Cited By
Dynamic interactive document generation method and device based on generative AI, computer equipment and computer readable storage medium
CN121072503A