Large model-based text automatic generation method and big data platform

By analyzing the topic-vocabulary distribution and vocabulary correlation of text data sets in various fields in the big data platform, and building attention weights based on the degree of difference, the transfer learning of text data sets in the source field to the target field is realized, solving the problem of performance decline in the field changes in the existing technology of text automatic generation large models, and improving the quality and accuracy of text generation.

CN119940361BActive Publication Date: 2025-06-27SHANDONG HAILIANXUN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510428998.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-06-27
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

The performance of existing text automatic generation large models declines when the data set or domain changes, and it is difficult to fully utilize the knowledge learned in the source domain pre-training stage in transfer learning, resulting in a decrease in the quality and accuracy of the target generated text.

Method used

By obtaining text data sets in various fields in the big data platform, performing topic-vocabulary distribution modeling, clustering vocabulary, analyzing the membership of vocabulary and clustering centers, determining the vocabulary relevance and topic relevance, and combining the degree of difference to construct attention weights, realizing the transfer learning of text data sets in the source field to the target field.

Benefits of technology

It improves the efficiency and efficiency of transfer learning, enhances the model's ability to understand vocabulary semantics, captures the correlation between the source field and the target field vocabulary, optimizes the transfer learning process, and improves the quality and accuracy of text generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940361B_ABST
    Figure CN119940361B_ABST
Patent Text Reader

Abstract

This application relates to the technical field of text data generation, and specifically relates to a method for automatic text generation based on a large model and a big data platform. The method includes: analyzing the membership degrees between each vocabulary in each field and different clustering centers, and determining the vocabulary association degree by combining the similarity of the vocabulary vectors of any two vocabularies between the source field and the target field; determining the topic association degree based on the vocabulary association degree between any two clustering centers between the source field and the target field, and combining the probability distributions of the any two clustering centers under different topics; determining the difference degree by analyzing the differences in the membership degree vectors of all clustering clusters between the source field and the target field, and migrating and learning the source field text data set to the target field in combination with the topic association degree. This application aims to optimize the way of transfer learning and improve the quality and accuracy of the target generated text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of text data generation, and specifically to a method for automatic text generation based on a large model and a big data platform. Background Art

[0002] Automatic text generation is to automatically generate output content according to a given input. Typical automatic text generation tasks include machine translation, text summarization, dialogue tasks, etc. The large model for automatic text generation trained by a deep neural network has excellent performance in some limited fields. However, when the dataset or field changes, the text generation performance of the large model often drops drastically. Moreover, the neural network model requires a large amount of data to support its training, and retraining the large model in a new field requires huge time and computing resources. Therefore, the high sensitivity of the large model for automatic text generation to dataset and field changes makes its performance in cross-domain text generation poor.

[0003] Transfer learning, as a commonly used machine learning method, applies the task parameters obtained from training to a new task, and transfers the knowledge learned by the model in a field to other tasks to improve the training efficiency and performance of the new task. A relatively mature method for automatic text generation in the industry is to fine-tune the downstream target field task in the pre-trained text generation model in the source field. In the prior art, the transfer learning method makes less use of the internal structure and semantic information correlation between text data in different fields. As the number of parameters of the pre-trained text generation model in the source field increases, it is difficult to align the text features of the pre-training task and the fine-tuning task during the transfer learning process. The text generation model cannot fully utilize the knowledge learned in the pre-training stage in the source field, thereby reducing the quality and accuracy of the target generated text. Summary of the Invention

[0004] In order to solve the above technical problems, the purpose of this application is to provide a method for automatic text generation based on a large model and a big data platform, and the specific technical solutions adopted are as follows:

[0005] In a first aspect, an embodiment of this application provides a method for automatic text generation based on a large model, and the method includes the following steps:

[0006] S1: In the big data platform, obtain text datasets in each field, and obtain the topic-lexical distribution of each field's text dataset. The probability of each word under all topics is used to form the word vector of each word, where each field includes a source field and a target field;

[0007] S2: Cluster all the words in the text datasets of each field, analyze the membership degrees of each word in each field with different cluster centers, determine the semantic relevance of each word in each field, and combine the similarity of the word vectors of any two words between the source field and the target field to determine the word correlation degree between any two words in the source field and the target field;

[0008] S3: Based on the word correlation degree between any two cluster centers in the source field and the target field, and combining the probability distributions of the any two cluster centers under different topics, determine the topic correlation degree of different topics between the source field and the target field;

[0009] S4: Based on the membership degree, obtain the membership degree vector of each cluster under each field; by analyzing the differences in the membership degree vectors of all clusters between the source field and the target field, determine the difference degree between the source field and the target field;

[0010] S5: Based on the topic correlation degree and the difference degree, determine the attention weights of different topics between the source field and the target field, and transfer the source field text dataset to the target field through transfer learning.

[0011] Preferably, the fuzzy C - means clustering algorithm is used to cluster all the words in the text datasets of each field, and the Euclidean distance between the word vector of the word and the word vector of the cluster center is used as the distance between the sample point and the cluster center in the objective function of the fuzzy C - means clustering algorithm.

[0012] Preferably, the method for determining the semantic relevance of each word in each field is as follows:

[0013] In each field, calculate the average membership degree of each word with all cluster centers. Among the membership degrees of each word with all cluster centers, obtain the membership degrees greater than the average membership degree, which are denoted as characteristic membership degrees. Take the sum of all characteristic membership degrees as the semantic relevance of each word in each field.

[0014] Preferably, the expression of the word correlation degree between any two words in the source field and the target field is: ; where represents the word correlation degree between the \(i\) - th word in the source field and the \(j\) - th word in the target field; represents the similarity between the word vector of the \(i\) - th word in the source field and the word vector of the \(j\) - th word in the target field; represents the difference in semantic relevance between the \(i\) - th word in the source field and the \(j\) - th word in the target field; represents a preset constant greater than 0.

[0015] Preferably, the method for determining the topic relevance between different topics in the source domain and the target domain is as follows:

[0016] Denote all the cluster centers as central words, and the topic relevance between topic m in the source domain and topic n in the target domain is expressed as: ; where represents the word relevance between the p-th central word in the source domain and the q-th central word in the target domain; represents the difference between the probability of the p-th central word under topic m in the source domain and the probability of the q-th central word under topic n in the target domain; represents the number of all central words in the target domain;

[0017] Preferably, the membership degree vector of each cluster under each domain consists of the membership degrees between all words in each cluster and the cluster center of the corresponding cluster in each domain.

[0018] Preferably, the expression for the difference degree between the source domain and the target domain is: ; where represents the difference degree between the source domain and the target domain; represents the difference between the membership degree vector of the h-th cluster in the source domain and the membership degree vector of the k-th cluster in the target domain; N represents the number of all clusters in the target domain; norm( ) represents the normalization function.

[0019] Preferably, the process for determining the attention weights of different topics between the source domain and the target domain is as follows:

[0020] The attention weight between topic m in the source domain and topic n in the target domain is expressed as: ; where represents the difference degree between the source domain and the target domain; represents the topic relevance between topic m in the source domain and topic n in the target domain; exp( ) represents the exponential function with the natural constant as the base;

[0021] Preferably, migrating the source domain text dataset to the target domain includes:

[0022] Constructing an attention mechanism layer in the feature alignment process of transfer learning through the attention weights of all topics between the source domain and the target domain;

[0023] The source domain text dataset is trained using the BERT model to obtain a pre-trained model; the source domain text dataset, the target domain text dataset, and the pre-trained model are used as the inputs of the transfer learning technique. Among them, the attention mechanism layer is used to align the features of the text datasets in the source domain and the target domain, and a text generation model adapted to the target domain is output.

[0024] In a second aspect, an embodiment of the present application also provides a large model-based text automated generation big data platform, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the above-mentioned large model-based text automated generation method are implemented.

[0025] The present application has at least the following beneficial effects:

[0026] By analyzing the similarity of the lexical vectors of any vocabulary between the source domain and the target domain, as well as the semantic relevance of the vocabulary, the present application constructs a word correlation degree, enhances the model's understanding ability of vocabulary semantics in the transfer learning process, can better capture the correlation between the source domain and the target domain vocabulary, and thus improves the efficiency of transfer learning; further, by comprehensively considering the word correlation degree between any two cluster centers between the source domain and the target domain, and combining the probability distributions of the any two cluster centers under different topics, a topic correlation degree is constructed, which helps the model understand the topic relationship between the source domain and the target domain, optimizes the strategy selection in the transfer learning process, and improves the text generation quality; further, by analyzing the differences in the membership vectors of all cluster clusters between the source domain and the target domain, a difference degree is constructed, which helps the model identify the differences between the source domain and the target domain, and thus adjust the transfer learning strategy and enhance the adaptability of the model in the target domain; finally, by comprehensively considering the topic correlation degree and the difference degree, an attention weight is constructed, which helps the model better align the text features of the source domain and the target domain in the transfer learning process, and improves the quality and accuracy of the target domain text generation. By analyzing the text semantic features, the present application optimizes the transfer learning method and improves the quality and accuracy of the target generated text. Description of the Drawings

[0027] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0028] Figure 1 It is a flowchart of the steps of a large model-based text automated generation method provided by an embodiment of the present application;

[0029] Figure 2 Schematic diagram of the attention weight extraction process provided by an embodiment of the present application. Detailed implementation manners

[0030] In order to further elaborate on the technical means and effects adopted by the present application to achieve the predetermined invention purpose, the following, in combination with the accompanying drawings and preferred embodiments, details the specific implementation manners, structures, features and effects of the text automation generation method and big data platform based on the large model proposed according to the present application. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.

[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs.

[0032] The following specifically describes the specific solutions of the text automation generation method and big data platform based on the large model provided by the present application with reference to the accompanying drawings.

[0033] Please refer to Figure 1 , which shows a flowchart of the steps of the text automation generation method based on the large model provided by an embodiment of the present application. The method includes the following steps:

[0034] S1: In the big data platform, obtain text data sets in each field, and obtain the topic-word distribution of each field's text data set. Combine the probabilities of each word under all topics to form a word vector for each word, where each field includes a source field and a target field.

[0035] In this embodiment, taking the big data platform as a carrier, using the text data sets stored in different fields therein, the training of the text generation model is realized through the big data processing framework and the deep learning framework. The topic modeling representation method of text data can effectively reduce the text feature dimension, thereby simplifying the subsequent processing and analysis process. However, the topic model has insufficient analysis of the semantic association degree between words in different topics, cannot handle the problems of "one meaning with multiple words" and "one word with multiple meanings", and it is difficult to completely transfer the knowledge obtained by training the model in the source field to the target field.

[0036] In this embodiment, text data sets in each field are obtained in the big data platform, and the LDA model is used to perform topic modeling on the text data sets in each field to obtain the topic-word distribution of each field's text data set. Among them, the topic-word distribution reflects the probability of a word under different topics. Therefore, further, the probabilities of each word under all topics are combined to form a word vector for each word. In this embodiment, each field includes a source field and a target field.

[0037] Among them, the principle of the LDA model is well-known technology, and the process of using the LDA model for topic modeling and obtaining the topic-word distribution will not be elaborated here.

[0038] S2: Cluster all the words in the text datasets of each field, analyze the membership degrees of each word in each field with different cluster centers, determine the semantic relevance of each word in each field, and combine the similarity of the word vectors of any two words between the source field and the target field to determine the word correlation degree between any two words in the source field and the target field.

[0039] Currently, when dealing with text data, transfer learning methods usually ignore the complex correlations between words in different topics. This kind of correlation is not only the direct connection between words, but also covers the semantic associations and interactions of words under multiple topics. The simple correlation of words, that is, the direct connection between words, although it can provide certain semantic information in some cases, in the text data summary, this kind of correlation is often not enough to capture the deep semantic structure and the subtle differences between topics. In fact, the word correlations in text data carry rich semantic information, and these information are crucial for improving the performance of transfer learning.

[0040] Therefore, in order to better mine the word correlations between the source field and the target field, the fuzzy C-means clustering (FCM) algorithm is used to cluster all the words in the text datasets of each field, obtaining multiple clusters and the membership degrees of each word in each field with different cluster centers;

[0041] Among them, the fuzzy C-means clustering algorithm is well-known technology. The Euclidean distance between the word vector of a word and the word vector of a cluster center is used as the distance between the sample point and the cluster center in the objective function of the fuzzy C-means clustering algorithm. The specific principle process of the fuzzy C-means clustering algorithm and the specific calculation process of the Euclidean distance will not be elaborated here.

[0042] Furthermore, in each field, calculate the mean value of the membership degrees of each word with all cluster centers. Among the membership degrees of each word with all cluster centers, obtain the membership degrees greater than the mean value of the membership degrees, denoted as characteristic membership degrees. Take the sum of all characteristic membership degrees as the semantic relevance of each word in each field.

[0043] Furthermore, by analyzing the similarity of the word vectors of any two words between the source field and the target field, and combining the semantic relevance, determine the word correlation degree between any two words in the source field and the target field, specifically:

[0044] The word correlation degree between the i-th word in the source field and the j-th word in the target field The expression of is as follows: represents the similarity between the lexical vector of the i-th word in the source domain and the lexical vector of the j-th word in the target domain; represents the difference in semantic relevance between the i-th word in the source domain and the j-th word in the target domain; represents a preset constant greater than 0, which is used to prevent the denominator from being 0. The value of is set artificially. In this embodiment,

[0045] According to the lexical association degree between any two words in the source domain and the target domain, it can be understood that if the similarity between the lexical vector of the i-th word in the source domain and the lexical vector of the j-th word in the target domain is greater, and the difference in semantic relevance between the i-th word in the source domain and the j-th word in the target domain is smaller, then the lexical association degree between the i-th word in the source domain and the j-th word in the target domain is greater, indicating that the association between words in the source domain and the target domain is greater; conversely, if the similarity between the lexical vector of the i-th word in the source domain and the lexical vector of the j-th word in the target domain is smaller, and the difference in semantic relevance between the i-th word in the source domain and the j-th word in the target domain is greater, then the lexical association degree between the i-th word in the source domain and the j-th word in the target domain is smaller, indicating that the association between words in the source domain and the target domain is smaller.

[0046] S3: Based on the lexical association degree between any two cluster centers in the source domain and the target domain, and combined with the probability distributions of the any two cluster centers under different themes, determine the theme association degrees of different themes between the source domain and the target domain.

[0047] For the sake of easy understanding, the words corresponding to the cluster centers are denoted as central words. The central words are the keywords in the context represented by the corresponding cluster clusters. The frequencies of different central words in the same theme represent the association between the theme and the contexts corresponding to different central words. The greater the lexical association degree between different central words across domains, the stronger the correlation between the corresponding pre-advances they represent. Therefore, taking the association degree between two central words as the weight, calculate the theme association degrees of different themes between the source domain and the target domain. Specifically:

[0048] The theme association degree between theme m in the source domain and theme n in the target domain The expression of

[0049] is as follows: It represents the lexical correlation degree between the p-th central word in the source domain and the q-th central word in the target domain; It represents the difference between the probability of the p-th central word under the theme m in the source domain and the probability of the q-th central word under the theme n in the target domain; It represents the number of all central words in the target domain; It represents a preset constant greater than 0, which is used to prevent the denominator from being 0. The value of is set artificially. In this embodiment, the value of is 0.01. On the premise of ensuring that the denominator is not 0 and does not overly affect the calculation result, the implementer can also set it according to the specific situation by himself / herself. This embodiment does not make special restrictions.

[0050] Among them, the source domain text dataset must be greater than or equal to the target domain text dataset.

[0051] Based on the theme correlation degree analysis between different themes of the source domain and the target domain: if the lexical correlation degree between the p-th central word in the source domain and the q-th central word in the target domain is larger, and the difference between the probability of the p-th central word under the theme m in the source domain and the probability of the q-th central word under the theme n in the target domain is smaller, then the theme correlation degree between the theme m in the source domain and the theme n in the target domain is larger, indicating that the theme relevance between the source domain and the target domain is larger and more conducive to text migration; on the contrary, if the lexical correlation degree between the p-th central word in the source domain and the q-th central word in the target domain is smaller, and the difference between the probability of the p-th central word under the theme m in the source domain and the probability of the q-th central word under the theme n in the target domain is larger, then the theme correlation degree between the theme m in the source domain and the theme n in the target domain is smaller, indicating that the theme relevance between the source domain and the target domain is smaller.

[0052] S4: Based on the membership degree, obtain the membership degree vector of each clustering cluster in each domain; by analyzing the differences of the membership degree vectors of all clustering clusters between the source domain and the target domain, determine the difference degree between the source domain and the target domain.

[0053] When there are large differences in text styles, vocabulary usage, etc. between the source domain and the target domain, it is difficult to align the theme features across domains, resulting in insufficient adaptability of the pre-trained text generation model in the source domain within the target domain. According to the lexical correlation degree and the theme correlation degree, an attention mechanism is added during the training process of transfer learning, and the attention mechanism automatically learns the alignment information of different themes across domains, enhances the model's utilization of the internal structure and semantic information of texts across domains, and further improves the text generation performance of the model. The specific process is as follows:

[0054] The membership degrees between all the words within each cluster in each domain and the cluster center of the corresponding cluster are used to form the membership degree vectors of each cluster under each domain.

[0055] The difference degree between the source domain and the target domain is expressed as: ; in the formula, represents the difference between the membership degree vector of the h-th cluster in the source domain and the membership degree vector of the k-th cluster in the target domain; N represents the number of all clusters in the target domain; norm( ) represents the normalization function.

[0056] It should be noted that there are many methods to measure the difference between vectors. In this embodiment, the DTW distance between the membership degree vector of the h-th cluster in the source domain and the membership degree vector of the k-th cluster in the target domain is used as the difference between the membership degree vector of the h-th cluster in the source domain and the membership degree vector of the k-th cluster in the target domain; in the actual application process, as other implementation manners, the implementer can also adopt other methods to measure the difference between vectors, such as Euclidean distance or Manhattan distance. Regarding the selection of the method to measure the difference between vectors, this embodiment does not make special restrictions.

[0057] Among them, the calculation method of the DTW distance is a well-known technology, and its specific calculation process will not be elaborated.

[0058] From the difference degree between the source domain and the target domain, it can be understood that if the difference between the membership degree vector of the h-th cluster in the source domain and the membership degree vector of the k-th cluster in the target domain is smaller, then the difference degree between the source domain and the target domain is smaller, indicating that the text style difference between the source domain and the target domain is smaller, the correlation degree is larger, and the accuracy of text migration is higher; on the contrary, if the difference between the membership degree vector of the h-th cluster in the source domain and the membership degree vector of the k-th cluster in the target domain is larger, then the difference degree between the source domain and the target domain is larger, indicating that the text style difference between the source domain and the target domain is larger, the correlation degree is smaller, and the accuracy of text migration is lower.

[0059] S5: Based on the topic correlation degree and the difference degree, determine the attention weights of different topics between the source domain and the target domain, and transfer the source domain text data set to the target domain through transfer learning.

[0060] The greater the topic correlation degree of different topics between different domains, the stronger the text information correlation of the corresponding topic in transfer learning, which helps to improve the adaptability of the model to the target domain. Therefore, a larger attention weight is set; the greater the text style difference formed by different context combinations across domains in the training set, the smaller the role of the model in transfer learning in the target domain, and the smaller the corresponding attention weight.

[0061] Therefore, according to the topic relevance and the difference degree, the attention weights of different topics between the source domain and the target domain are determined, and the specific process is as follows:

[0062] The attention weight between topic m in the source domain and topic n in the target domain is expressed as: ; in the formula, represents the difference degree between the source domain and the target domain; represents the topic relevance between topic m in the source domain and topic n in the target domain; exp( ) represents the exponential function with the natural constant as the base; represents the normalization function.

[0063] It can be understood from the attention weights of different topics between the source domain and the target domain that if the difference degree between the source domain and the target domain is smaller, and the topic relevance between topic m in the source domain and topic n in the target domain is larger, it indicates that the text information relevance is stronger in the transfer learning process, which helps to improve the adaptability of the model to the target domain. Therefore, the attention weight between topic m in the source domain and topic n in the target domain is larger, and a larger attention weight needs to be set to improve the accuracy of transfer learning; on the contrary, if the difference degree between the source domain and the target domain is larger, and the topic relevance between topic m in the source domain and topic n in the target domain is smaller, it indicates that the text information relevance is weaker in the transfer learning process, and the text style difference is larger. Therefore, the attention weight between topic m in the source domain and topic n in the target domain is smaller.

[0064] Preferably, the schematic diagram of the attention weight extraction process provided in this embodiment is as Figure 2 shown.

[0065] Based on the topic relevance and the difference degree, the attention weight is obtained. Further, an attention mechanism layer in the feature alignment process of transfer learning is constructed through the attention weights of all topics between the source domain and the target domain.

[0066] The source domain text dataset is trained using the BERT model to obtain a pre-trained model; the source domain text dataset, the target domain text dataset, and the pre-trained model are used as the input of the transfer learning technology. Among them, the attention mechanism layer is used to align the features of the text datasets in the source domain and the target domain, and a text generation model adapted to the target domain is output.

[0067] Among them, the attention mechanism layer in the feature alignment process of transfer learning is constructed through the attention weights of all topics between the source domain and the target domain. The BERT model and the transfer learning technology are all well-known technologies, and their specific principle processes will not be elaborated here.

[0068] So far, in this embodiment, by analyzing the semantic relevance between the source domain and the target domain, an attention mechanism is added during the model transfer learning process to ensure the accurate alignment between the text features of the source domain and the target domain, enabling the model to learn knowledge within the source domain from multiple perspectives, improving the effect of transfer learning, and further enhancing the performance of the text generation of the text automatic generation model within the target domain.

[0069] Based on the same inventive concept as the above method, an embodiment of the present application also provides a big data platform for text automatic generation based on a large model, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of any one of the above methods for text automatic generation based on a large model.

[0070] It should be noted that the above sequence of embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. And the above specific embodiments of this specification have been described. Additionally, the processes depicted in the drawings do not necessarily require the specific order or consecutive order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0071] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments.

[0072] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present application shall be included within the protection scope of the present application.

Claims

1. A method for automatic text generation based on a large model, characterized in that: The method comprises the following steps: S1: In the big data platform, obtain text datasets in various fields and the topic-vocabulary distribution of the text datasets in various fields, and use the probability of each word under all topics to form the vocabulary vector of each word, where each field includes a source field and a target field; S2: Cluster all words in the text data set of each field. In each field, calculate the mean membership between each word and all cluster centers. Among the membership between each word and all cluster centers, obtain the membership greater than the mean membership, record it as the feature membership, and take the cumulative sum of all feature memberships as the semantic relevance of each word in each field. Combined with the similarity of the vocabulary vectors of any two words between the source field and the target field, determine the vocabulary relevance between any two words between the source field and the target field. S3: Based on the vocabulary association between any two cluster centers between the source domain and the target domain, and in combination with the probability distribution of the any two cluster centers under different topics, determine the topic association between different topics between the source domain and the target domain; S4: Based on the membership, obtaining the membership vector of each cluster in each field; determining the difference between the source field and the target field by analyzing the difference of the membership vectors of all clusters between the source field and the target field; S5: Based on the topic relevance and the difference, determine the attention weights of different topics between the source domain and the target domain, and transfer the source domain text dataset to the target domain.

2. The method for automatically generating text based on a large model as claimed in claim 1, characterized in that: The fuzzy C-means clustering algorithm is used to cluster all vocabulary of text data sets in various fields, and the Euclidean distance between the vocabulary vector of the vocabulary and the vocabulary vector of the cluster center is used as the distance between the sample point and the cluster center in the objective function of the fuzzy C-means clustering algorithm.

3. The method for automatically generating text based on a large model as claimed in claim 1, characterized in that: The expression of the vocabulary association between any two words in the source domain and the target domain is: ; In the formula, represents the lexical association between the i-th word in the source domain and the j-th word in the target domain; Represents the similarity between the vocabulary vector of the i-th word in the source domain and the vocabulary vector of the j-th word in the target domain; represents the difference in semantic relevance between the i-th word in the source domain and the j-th word in the target domain; Indicates a preset constant greater than 0.

4. The method for automatically generating text based on a large model as claimed in claim 1, characterized in that: The method for determining the subject relevance of different subjects between the source domain and the target domain is as follows: All cluster centers are recorded as central words, and the topic relevance between topic m in the source domain and topic n in the target domain is The expression is: ; In the formula, represents the lexical relevance between the p-th central word in the source domain and the q-th central word in the target domain; represents the difference between the probability of the pth central word under topic m in the source domain and the probability of the qth central word under topic n in the target domain; represents the number of all central words in the target domain; Indicates a preset constant greater than 0.

5. The method for automatically generating text based on a large model as claimed in claim 1, characterized in that: The membership vector of each cluster in each field is composed of the membership between all words in each cluster in each field and the cluster center of the corresponding cluster.

6. The method for automatically generating text based on a large model as claimed in claim 1, characterized in that: The expression of the difference between the source domain and the target domain is: ; In the formula, Indicates the difference between the source domain and the target domain; represents the difference between the membership vector of the hth cluster in the source domain and the membership vector of the kth cluster in the target domain; N represents the number of all clusters in the target domain; norm() represents the normalization function.

7. The method for automatically generating text based on a large model as claimed in claim 1, characterized in that: The process of determining the attention weights of different topics between the source domain and the target domain is as follows: The attention weight between topic m in the source domain and topic n in the target domain The expression is: ; In the formula, Indicates the difference between the source domain and the target domain; represents the topic relevance between topic m in the source domain and topic n in the target domain; exp( ) represents an exponential function with a natural constant as the base; Represents the normalization function.

8. The method for automatically generating text based on a large model as claimed in claim 1, characterized in that: The method of transferring the source domain text dataset to the target domain includes: Constructing the attention mechanism layer in the feature alignment process in transfer learning through the attention weights of all topics between the source domain and the target domain; The BERT model is used to train the source domain text dataset to obtain a pre-trained model; the source domain text dataset, the target domain text dataset and the pre-trained model are used as inputs of the transfer learning technology, wherein the attention mechanism layer is used to align the features of the text datasets in the source domain and the target domain, and the text generation model adapted to the target domain is output.

9. A big data platform for automatic text generation based on a big model, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the large model-based automatic text generation method as described in any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Comment emotion classification method and system based on deep hybrid model transfer learning

    CN109271522A

  • Field-adaptive deep knowledge tracking and personalized exercise recommendation method

    CN111444432A