Scenarized AIGC content generation method and system
Through the collaborative work of the bootstrap and the artificial intelligence content generation engine, the scene classification model and preset database are used to generate an update prompt word sequence that matches the target scenario, solving the problem of insufficient scenario in the existing AIGC method and achieving efficient and personalized content generation.
Patent Information
- Application Number
- CN202510362424.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-11
AI Technical Summary
The existing AIGC content generation methods lack scenario-based design and are difficult to meet the personalized needs in specific scenarios. Users need to adjust the prompt words repeatedly, which is inefficient.
Through the collaborative work of the guide and the artificial intelligence content generation engine, the scene classification model and preset database are used to generate an update prompt word sequence that matches the target scene, and optimize the lexicon for user interaction to achieve scene-based generation of content.
It improves the pertinence and practicality of the content, reduces the user's expression ability requirements, improves the generation efficiency and accuracy, and enhances the richness and diversity of the content.
Smart Images

Figure CN120297411A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to data processing technologies, and particularly to a scenario-based AIGC content generation method and system. Background Art
[0002] In the current era of rapid informatization development, content creation has become an indispensable part of various industries. Whether in the fields of office work, advertising, media, education, or entertainment, a large amount of high-quality content is required to attract and meet user needs. With the continuous progress of artificial intelligence technologies, especially the rapid development of technologies such as natural language processing and generative adversarial networks, artificial intelligence generated content (AIGC) has gradually become a new hot spot in the content creation field. The AIGC technology can automatically generate various forms of content such as text, images, audio, or videos based on given prompts or topics, greatly improving the efficiency and diversity of content creation.
[0003] However, existing AIGC content generation methods are often designed for generalization and lack pertinence. That is, when the scenario is not clearly stated in the prompt statement, they usually do not consider the specific scenario of content generation, but create content based on general prompt information. Although this "one-size-fits-all" approach can generate content, it is often difficult to meet the personalized needs in specific scenarios. If more matching content needs to be generated, users are required to repeatedly adjust and try the prompt words or statements, which has high requirements for users' expression ability and usually requires multiple rounds of interaction, resulting in low overall efficiency. Summary of the Invention
[0004] The purpose of the present invention is to provide a scenario-based AIGC content generation method and system, which can more accurately identify the user's intention and determine the target scenario that matches it, so as to ensure that the generated content better meets the needs of specific scenarios and improve the pertinence and practicality of the content.
[0005] The purpose of the present invention is achieved through the following technical solutions:
[0006] A scenario-based AIGC content generation method, characterized in that it uses a scenario-based content generation system to complete AIGC content generation. The scenario-based content generation system includes a guide and an artificial intelligence content generation engine; specifically, it includes the following steps:
[0007] S101 Obtain the original prompt information input by the user through the guide, perform word segmentation processing on the original prompt information, and generate an original prompt word sequence;
[0008] The S102 guide inputs the original prompt sequence into a preset scenario classification model to determine the predicted target scenario corresponding to the original prompt information. The preset scenario classification model is a classifier generated based on a deep learning algorithm;
[0009] The S103 guide generates an updated prompt sequence according to the predicted target scenario and the original prompt sequence. The updated prompt sequence includes an original prompt sequence and an extended prompt sequence. The extended prompt sequence is a prompt sequence obtained by extending the original prompt sequence in the direction of the predicted target scenario;
[0010] The S104 guide inputs the original prompt information and the updated prompt sequence into an artificial intelligence content generation engine to generate artificial intelligence-generated content corresponding to the prompt information.
[0011] By parsing and classifying the original prompt information through a preset scenario classification model, the user intention can be recognized, and the target scenario that matches it can be determined, so as to ensure that the generated content better meets the requirements of a specific scenario, improving the pertinence and practicality of the content. Among them, the original prompt sequence is extended according to the predicted target scenario to generate an updated prompt sequence containing more relevant words and expressions, thereby improving the richness of the subsequent generated content. In addition, through the collaborative work of the guide and the artificial intelligence content generation engine, the process from user input to content generation is realized, reducing the usage threshold of users, avoiding the need for repeated adjustment of prompt words, thus improving the generation efficiency and ensuring the high quality and personalization of the content.
[0012] Further, the guide generating an updated prompt sequence according to the predicted target scenario and the original prompt sequence includes:
[0013] The guide obtains a target scenario extension word library from a preset database according to the predicted target scenario;
[0014] The guide determines the similarity between each original prompt word in the original prompt sequence and each extension word in the target scenario extension word library, so as to add the extension words with a similarity greater than a preset similarity threshold to the extended prompt sequence;
[0015] The guide generates the updated prompt sequence according to the original prompt sequence and the extended prompt sequence.
[0016] By obtaining a target scenario expansion vocabulary library that matches the predicted target scenario from a preset database, the bootstrapper can further expand the vocabulary closely related to a specific scenario, thereby ensuring that the generated updated prompt sequence is highly relevant to the original prompt information and closely centered around the target scenario, thus improving the relevance and pertinence of the generated content. Among them, through similarity calculation, the bootstrapper adds the expansion words with a similarity greater than the preset threshold in the target scenario expansion vocabulary library to the expanded prompt sequence, which not only enriches the vocabulary selection of the content but also increases the diversity and expandability of the content. In addition, by using deep learning algorithms and a preset database, intelligent parsing and expansion of the original prompt information can be achieved, that is, the bootstrapper can automatically complete complex tasks from scenario classification to vocabulary expansion without manual intervention, thereby improving the intelligent level and efficiency of content generation. And since the generated updated prompt sequence is more in line with the user's intention and scenario requirements, the generated content is also closer to the user's expectations.
[0017] Further, the scenario-based content generation system includes a cluster of terminal devices and a server. Each terminal device in the cluster of terminal devices is communicatively connected to the server through a local area network. The server is communicatively connected to the artificial intelligence content generation engine provided on a wide area network. The preset database is provided in the server, and the bootstrapper is provided in the terminal device and / or the server;
[0018] Correspondingly, the bootstrapper inputs the original prompt information and the updated prompt sequence into a preset artificial intelligence content generation engine, including:
[0019] The bootstrapper inputs the original prompt information and the updated prompt sequence into the artificial intelligence content generation engine through the wide area network;
[0020] Correspondingly, after generating the artificial intelligence-generated content corresponding to the prompt information, it further includes:
[0021] The artificial intelligence content generation engine sends the artificial intelligence-generated content to the server for display on the terminal device.
[0022] The local area network communication connection between the terminal device and the server, as well as the wide area network communication between the server and the artificial intelligence content generation engine, can not only ensure the security of internal data transmission but also guarantee access to external artificial intelligence content generation engines. The booter inputs the original prompt information and the updated prompt word sequence into the artificial intelligence content generation engine through the wide area network, achieving fast data transfer and processing. At the same time, the generated content is transmitted back to the terminal device through the server for display, ensuring the real-time and accuracy of the content. In addition, the design of the system architecture fully considers the effective utilization of resources and load balancing. The terminal device cluster is responsible for user interaction and data processing, and the server is responsible for centralized data storage and distribution, thus ensuring the security of the internal network.
[0023] Optionally, the booter generates the updated prompt word sequence according to the original prompt word sequence and the extended prompt word sequence, including:
[0024] The booter displays a prompt word network diagram through the terminal device, where the prompt word network diagram includes the original prompt word node sequence corresponding to the original prompt word sequence, the extended prompt word node sequence corresponding to the extended prompt word sequence, and a feature edge sequence for establishing the mapping relationship between each original prompt word node in the original prompt word node sequence and each extended prompt word node in the extended prompt word node sequence;
[0025] Obtain a weight adjustment instruction for a target feature edge in the feature edge sequence through the booter, and only determine the extended prompt word corresponding to the target feature edge with a weight greater than the first preset weight threshold after adjustment as the target extended prompt word;
[0026] Generate the updated prompt word sequence according to the original prompt word sequence and all the target extended prompt words.
[0027] By displaying the prompt word network diagram on the terminal device, users can intuitively see the relationship network between the original prompt words and the extended prompt words. This visual display method enhances the interactivity between users and the content generation system, enabling users to more clearly understand the logic and process of content generation. At the same time, users can adjust the weights of the feature edges according to their own needs, thus realizing personalized customization of the generated content and improving the pertinence and practicality of the content. Then, by obtaining the user's adjustment instruction for the feature edge weight, the system can more accurately understand the user's intentions and needs. The system only includes the extended prompt words corresponding to the feature edges with weights greater than the preset threshold after weight adjustment in the updated prompt word sequence. This screening mechanism ensures that the generated content is more in line with the actual needs of users and improves the quality and accuracy of the content.
[0028] Further, after obtaining the weight adjustment instruction for the target feature edge in the feature edge sequence through the guide, the following steps are also included:
[0029] The guide adds the original prompt words corresponding to the target feature edges whose adjusted weights are greater than the second preset weight threshold to the target scene extension word library.
[0030] By adding the original prompt words with higher weights in the user feedback to the target scene extension word library, the system can optimize and expand its word library content in real time and dynamically. This process ensures that the word library can keep up with user needs and language development trends, improving the timeliness and relevance of the system. As the word library continues to be enriched and optimized, the system can more accurately understand and match the user's intentions and scene requirements when generating AIGC content, which helps to improve the intelligence level of content generation and make the generated content more in line with the user's expectations and requirements. In addition, by continuously absorbing user feedback to optimize itself, the system can promote the expansion and optimization of the word library. When facing diverse user needs and complex language environments, the system can exhibit higher robustness and adaptability.
[0031] Further, the preset scene classification model includes a prompt word input layer, a convolutional neural network layer, a long short-term memory network layer, and a scene output layer;
[0032] The prompt word input layer is used to input the original prompt word sequence and convert the original prompt word sequence into an original prompt word vector sequence to transfer the original prompt word vector sequence to the convolutional neural network layer;
[0033] The convolutional neural network layer is used to extract features from the original prompt word vector sequence to generate an original prompt word feature sequence and transfer the original prompt word feature sequence to the long short-term memory network layer;
[0034] The long short-term memory network layer is used to generate a hidden state sequence according to the original prompt word feature sequence and transfer the hidden state sequence to the scene output layer;
[0035] The scene output layer inputs the hidden state sequence into a preset classifier to output a predicted candidate scene sequence, the predicted candidate scene sequence includes at least one candidate scene sequence, and the predicted candidate scene sequence includes the predicted target scene.
[0036] The prompt input layer can convert the original prompt word sequence into a vector sequence, providing a standardized data format for subsequent processing and improving the efficiency of data processing. The convolutional neural network layer can capture and strengthen key information while filtering out irrelevant details by extracting features from the original prompt word vector sequence. The long short-term memory network layer effectively retains the temporal information of the data by generating a hidden state sequence, enabling the model to understand and process complex context relationships. The scenario output layer uses a preset classifier to process the hidden state sequence and output a predicted candidate scenario sequence, including the predicted target scenario. This process ensures the accuracy and reliability of scenario classification and provides strong support for subsequent content generation. By integrating the functions of the above layers, the preset scenario classification model can analyze the original prompt information, thereby guiding the content generation engine to generate content that better meets the target scenario and user needs.
[0037] Further, the convolutional neural network layer is used to extract features from the original prompt word vector sequence to generate an original prompt word feature sequence, including:
[0038] The convolutional neural network layer uses Equation 1 and generates the original prompt word feature sequence Y according to the original prompt word vector sequence X, where Equation 1 is:
[0039] Y = σ(W * X + b)
[0040] where σ is the activation function, W is the weight matrix of the convolutional kernel, * is the convolution operation, and b is the bias term;
[0041] Correspondingly, the long short-term memory network layer is used to generate a hidden state sequence according to the original prompt word feature sequence, including:
[0042] The long short-term memory network layer uses Equation 2 and generates the hidden state h at the current time step in the hidden state sequence H according to the original prompt word feature sequence Y t , where Equation 2 is:
[0043] h t = LSTM(h t-1 , Y)
[0044] where LSTM is the operation of the long short-term memory network layer, and h t-1 is the hidden state at the previous time step in the hidden state sequence H.
[0045] The convolutional neural network layer extracts features from the original prompt word vector sequence through Formula 1, thereby effectively extracting key features from the input data and forming a more representative original prompt word feature sequence. The long short-term memory network layer uses the LSTM mechanism through Formula 2 to combine the original prompt word feature sequence with the hidden state of the previous time step to generate the hidden state of the current time step, thereby effectively capturing and utilizing the context information in the input data, enabling the model to maintain memory of previous information when processing sequence data, and thus improving the processing ability for complex sequence data.
[0046] A scenario-based content generation system includes: a booter and an artificial intelligence content generation engine;
[0047] Obtain the original prompt information input by the user through the booter, and perform word segmentation processing on the original prompt information to generate an original prompt word sequence;
[0048] The booter inputs the original prompt word sequence into a preset scenario classification model to determine the predicted target scenario corresponding to the original prompt information, and the preset scenario classification model is a classifier generated based on a deep learning algorithm;
[0049] The booter generates an updated prompt word sequence according to the predicted target scenario and the original prompt word sequence, and the updated prompt word sequence includes the original prompt word sequence and an extended prompt word sequence, and the extended prompt word sequence is a prompt word sequence obtained by expanding the original prompt word sequence in the direction of the predicted target scenario;
[0050] The booter inputs the original prompt information and the updated prompt word sequence into the artificial intelligence content generation engine to generate the artificial intelligence generated content corresponding to the prompt information.
[0051] The electronic device of the present invention includes:
[0052] A processor; and,
[0053] A memory for storing executable instructions of the processor;
[0054] Wherein, the processor is configured to execute any possible method described in the first aspect by executing the executable instructions.
[0055] Computer-executable instructions are stored in the computer-readable storage medium of the present invention, and when the computer-executable instructions are executed by a processor, they are used to implement any possible method described in the first aspect.
[0056] The beneficial effects of the present invention are:
[0057] Obtain the original prompt information input by the user through the guide, perform word segmentation on the original prompt information to generate the original prompt word sequence, then input the original prompt word sequence into the preset scenario classification model to determine the predicted target scenario corresponding to the original prompt information, thereby generating an updated prompt word sequence according to the predicted target scenario and the original prompt word sequence, and then input the original prompt information and the updated prompt word sequence into the artificial intelligence content generation engine to generate the artificial intelligence generated content corresponding to the prompt information, thereby performing targeted expansion of the input original prompt information in a scenario-based manner to more accurately identify the user's intention, determine the target scenario that matches it, ensure that the generated content better meets the requirements of a specific scenario, improve the pertinence and practicality of the content, and reduce the requirements for the user's expression ability, reduce the number of interaction rounds, and effectively improve the generation efficiency and accuracy of the artificial intelligence generated content. Brief Description of the Drawings
[0058] Figure 1 is a schematic flowchart of a scenario-based AIGC content generation method according to an exemplary embodiment of the present invention;
[0059] Figure 2 is a schematic flowchart of a scenario-based AIGC content generation method according to another exemplary embodiment of the present invention;
[0060] Figure 3 is a schematic structural diagram of a scenario-based content generation system according to an exemplary embodiment of the present invention;
[0061] Figure 4 is a schematic structural diagram of an electronic device according to an exemplary embodiment of the present invention. Detailed Description of the Embodiments
[0062] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0063] To solve the above problems, in the embodiments provided by the present invention, first, obtain the original prompt information input by the user through the guide, and use word segmentation processing technology to convert it into the original prompt word sequence. Subsequently, input these prompt word sequences into the preset scenario classification model. This scenario classification model is constructed based on deep learning algorithms and can automatically learn and identify the characteristics of different scenarios, thereby accurately classifying the input information into the corresponding scenarios.
[0064] After identifying the predicted target scenario corresponding to the original prompt information, this embodiment further expands the original prompt word sequence according to this scenario. Specifically, the system retrieves an extended vocabulary related to the target scenario from a preset database, and by calculating the similarity between the original prompt words and the extended words, filters out the extended vocabulary that is highly relevant to the original prompt information and suitable for the target scenario. These extended vocabulary will be merged with the original prompt word sequence to form an updated prompt word sequence.
[0065] In addition, to enhance the flexibility of the system and the user experience, this embodiment also introduces a display method of a prompt word network graph. The relationship network between the original prompt words and the extended prompt words is displayed to the user through a terminal device, and the user is allowed to adjust the weights of the feature edges according to actual needs. The system will update the prompt word sequence in real time according to the user's adjustment results, so as to ensure that the generated content better meets the user's expectations and needs.
[0066] To continuously improve the accuracy and practicality of the system, this embodiment also designs a dynamic vocabulary optimization mechanism. Specifically, the system will monitor and analyze the user's feedback information in real time, especially the user's weight adjustment instructions for the extended prompt words. When the weight of an extended prompt word is adjusted by the user to exceed a preset threshold, the system will regard it as an important word and automatically add it to the extended vocabulary of the target scenario. This mechanism not only helps to improve the vocabulary richness and accuracy of the system, but also enables the system to better adapt to the needs and preferences of different users.
[0067] After generating the updated prompt word sequence, the system inputs it together with the original prompt information into the artificial intelligence content generation engine. Based on advanced natural language processing and generation technologies, this engine can automatically generate high-quality content that is highly relevant to the input information and meets the requirements of the target scenario. To improve the efficiency and accuracy of content generation, this embodiment also adopts a system architecture that combines a terminal device cluster, a server, and an artificial intelligence content generation engine. The terminal device is responsible for interacting with the user and collecting input information; the server is responsible for processing and analyzing this information and transmitting it to the artificial intelligence content generation engine for content generation; the finally generated content is then transmitted back to the terminal device through the server for the user to view and use.
[0068] Figure 1 is a schematic flowchart of a scenario-based AIGC content generation method according to an exemplary embodiment of the present invention. As Figure 1 shown, the scenario-based AIGC content generation method provided in this embodiment includes:
[0069] S101. Obtain the original prompt information input by the user through a booter, and perform word segmentation processing on the original prompt information to generate an original prompt word sequence.
[0070] In this step, the booter obtains the original prompt information input by the user. The booter performs word segmentation on the original prompt information, splitting the continuous text into independent lexical units to generate the original prompt word sequence. Word segmentation is a basic task in natural language processing and helps with subsequent understanding and processing of the prompt information.
[0071] Specifically, to provide a good user experience, the user can input the required original prompt information through the input interface of a terminal device (such as a computer, mobile phone, tablet, etc.). Among them, the interface can provide various input methods such as text boxes and voice input buttons to meet the needs of different users. For example, for users with poor eyesight, a voice input function can be provided; for users who prefer keyboard input, a clear text box is provided. The booter is responsible for receiving the original prompt information input by the user through the interface. These information will be temporarily stored in the local cache or memory for subsequent processing.
[0072] The booter is built-in with advanced Chinese word segmentation algorithms, such as dictionary-based word segmentation methods, statistics-based word segmentation methods, or a combination of both. For Chinese text, word segmentation is a basic step in processing natural language, which can split continuous text into individual independent lexical units.
[0073] To improve the accuracy of word segmentation, the word segmentation algorithm in the booter usually relies on a pre-built dictionary. This dictionary contains a large number of common words and phrases. As the user input continues to increase and change, the booter also has the ability to dynamically update the dictionary. For example, when encountering newly emerged internet buzzwords or professional terms, the booter can automatically add them to the dictionary so that they can be correctly recognized in future word segmentation processes.
[0074] While performing word segmentation, the booter can also perform part-of-speech tagging (such as nouns, verbs, adjectives, etc.) and named entity recognition (such as person names, place names, organization names, etc.) on each word. Although these information are not directly used to generate the original prompt word sequence, they can provide valuable references for subsequent content generation. After word segmentation processing, the booter arranges the obtained lexical units in the order of the input text to form the original prompt word sequence.
[0075] S102. The booter inputs the original prompt word sequence into a preset scenario classification model to determine the predicted target scenario corresponding to the original prompt information.
[0076] In this step, the booter inputs the original prompt word sequence into a preset scenario classification model to determine the predicted target scenario corresponding to the original prompt information. The preset scenario classification model is a classifier generated based on deep learning algorithms.
[0077] Specifically, the booter can input the original prompt sequence into a preset scene classification model. This scene classification model is a classifier generated based on deep learning algorithms and can identify and classify different scenes. Through model processing, the predicted target scene corresponding to the original prompt information is determined.
[0078] Furthermore, the preset scene classification model can be constructed using deep learning algorithms and has a multi-layer neural network structure. The model includes multiple parts such as an input layer, a hidden layer, and an output layer. The input layer is responsible for receiving the original prompt sequence. The hidden layer extracts features through multiple non-linear transformations. The output layer is responsible for outputting the scene classification result. When processing the input data, the model first extracts features from the original prompt sequence. This process uses techniques such as convolutional neural networks to capture the semantic relationships and context information between words and converts the text data into high-dimensional feature vectors. The extracted feature vectors are then sent into the classifier for scene classification. The classifier performs linear transformations and non-linear activations on the feature vectors based on pre-trained weight and bias parameters and finally outputs the predicted target scene label. This label represents the scene category corresponding to the original prompt information.
[0079] To ensure the accuracy of scene classification, the model usually attaches a probability value when outputting the scene label. This probability value represents the confidence level of the model in the classification result. In practical applications, a probability threshold can be set, and only when the probability of the predicted scene exceeds this threshold is it regarded as a valid predicted target scene. In some cases, the original prompt information may correspond to multiple scenes simultaneously. To handle this situation, the model can output a list of scene probability distributions and sort them according to the probability size. Then, the top N scenes with the highest probability can be selected as candidate target scenes for subsequent processing according to actual needs.
[0080] To improve the accuracy and generalization ability of the scene classification model, a user feedback mechanism can also be introduced. When the user expresses satisfaction with the generated content or proposes modification opinions, these feedback messages can be used to adjust the weight and bias parameters of the model, thereby optimizing the model performance. With the continuous addition of new data and the continuous training of the model, the scene classification model can gradually learn the characteristics and laws of more scenes, thereby continuously improving its classification ability and generalization ability.
[0081] For the training of the above model, a large number of labeled scene data can be collected first. These data should cover a variety of different scene categories. The collected data is labeled to clarify the scene category to which each data belongs. This usually requires manual completion to ensure the accuracy of the labeling.
[0082] Then, randomly initialize the weight and bias parameters of the model, or use a pre-trained model for initialization to accelerate convergence. Select an appropriate loss function (such as cross-entropy loss) and optimizer (such as Adam optimizer) to calculate the prediction error and update the model parameters during training. Input the preprocessed data into the model for multiple rounds of iterative training. In each iteration, perform forward propagation to calculate the prediction results, and perform backward propagation to calculate the gradients and update the model parameters. Use the validation set and test set to evaluate the performance of the model. Evaluate the classification effect of the model in different scenarios by calculating metrics such as accuracy and recall.
[0083] Next, adjust hyperparameters such as the learning rate and batch size according to the performance of the validation set to optimize the training effect of the model. Moreover, increase the diversity of the data through data processing operations to improve the generalization ability of the model. Combine the prediction results of multiple models to improve the overall performance, such as using ensemble learning methods like Bagging and Boosting.
[0084] After training is completed, deploy the model to the actual application scenario. Before deployment, the model can be compressed and optimized to reduce the volume of the model and improve the inference speed. In actual applications, according to the original prompt information input by the user, use the trained scenario classification model to determine the predicted target scenario, and generate the corresponding content accordingly.
[0085] In a possible implementation, the above-mentioned preset scenario classification model may include a prompt word input layer, a convolutional neural network layer, a long short-term memory network layer, and a scenario output layer. Among them, the prompt word input layer is used to input the original prompt word sequence and convert the original prompt word sequence into an original prompt word vector sequence to transfer the original prompt word vector sequence to the convolutional neural network layer. The convolutional neural network layer is used to extract features from the original prompt word vector sequence to generate an original prompt word feature sequence and transfer the original prompt word feature sequence to the long short-term memory network layer. The long short-term memory network layer is used to generate a hidden state sequence based on the original prompt word feature sequence and transfer the hidden state sequence to the scenario output layer. The scenario output layer inputs the hidden state sequence into the preset classifier to output a predicted candidate scenario sequence, and the predicted candidate scenario sequence includes at least one candidate scenario sequence, and the predicted candidate scenario sequence includes the predicted target scenario.
[0086] In the above solution, the prompt input layer can convert the original prompt word sequence into a vector sequence, providing a standardized data format for subsequent processing and improving the efficiency of data processing. The convolutional neural network layer can capture and strengthen key information while filtering out unimportant details by extracting features from the original prompt word vector sequence. The long short-term memory network layer effectively preserves the temporal information of the data by generating a hidden state sequence, enabling the model to understand and process complex context relationships. The scenario output layer processes the hidden state sequence using a preset classifier to output a predicted candidate scenario sequence, which includes the predicted target scenario. This process ensures the accuracy and reliability of scenario classification and provides strong support for subsequent content generation. By integrating the functions of the above layers, the preset scenario classification model can analyze the original prompt information, thereby guiding the content generation engine to generate content that better meets the target scenario and user needs.
[0087] Further, for the convolutional neural network layer to extract features from the original prompt word vector sequence to generate an original prompt word feature sequence, it may include:
[0088] The convolutional neural network layer uses Equation 1 and generates an original prompt word feature sequence Y based on the original prompt word vector sequence X, where Equation 1 is:
[0089] Y = σ(W * X + b)
[0090] where σ is the activation function, W is the weight matrix of the convolutional kernel, * is the convolution operation, and b is the bias term.
[0091] Correspondingly, for the long short-term memory network layer to generate a hidden state sequence based on the original prompt word feature sequence, it may include:
[0092] The long short-term memory network layer uses Equation 2 and generates the hidden state h at the current time step in the hidden state sequence H based on the original prompt word feature sequence Y t , where Equation 2 is:
[0093] h t = LSTM(h t-1 , Y)
[0094] where LSTM is the operation of the long short-term memory network layer, and h t-1 is the hidden state at the previous time step in the hidden state sequence H.
[0095] Correspondingly, for the scenario output layer to input the hidden state sequence into the preset classifier to output a predicted candidate scenario sequence, it may include:
[0096] The scenario output layer uses Equation 3 and inputs the hidden state sequence H into a preset classifier to output a predicted candidate scenario sequence D, where Equation 3 is:
[0097]
[0098] where softmax() is a preset normalization function used to convert the output into a probability distribution, W s is the weight matrix of the preset classifier, and b s is the bias term of the preset classifier, h T is the last hidden state in the hidden state sequence H, and P(D) is the classification probability value of the predicted candidate scenario sequence D, is the preset probability threshold;
[0099] The scenario output layer uses Equation 4 and determines the predicted target scenario d based on the predicted candidate scenario sequence D, where the formula
[0100] 4 is:
[0101]
[0102] where N is the number of predicted candidate scenarios in the predicted candidate scenario sequence D, w si is the weight vector corresponding to the i-th candidate scenario in the preset classifier, and b si is the bias term corresponding to the i-th candidate scenario in the preset classifier.
[0103] In the above solution, the convolutional neural network layer extracts features from the original prompt word vector sequence through Equation 1, thereby effectively extracting key features from the input data and forming a more representative original prompt word feature sequence. The long short-term memory network layer uses the LSTM mechanism through Equation 2 to combine the original prompt word feature sequence with the hidden state of the previous time step to generate the hidden state of the current time step, thereby effectively capturing and utilizing the context information in the input data, enabling the model to maintain memory of previous information when processing sequence data, and thus improving the processing ability for complex sequence data.
[0104] In addition, the scenario output layer inputs the hidden state sequence into the preset classifier through Equation 3 and Equation 4 for scenario classification and outputs a conversion into a probability distribution, enabling the model to output multiple possible scenario predictions and determining the final predicted target scenario by comparing the probability magnitudes, thereby not only improving the accuracy of scenario classification but also enhancing the reliability and accuracy of the classification results by setting a probability threshold and comparing probability magnitudes.
[0105] Specifically, the above formula 3 represents the decision function used by the scene output layer during scene classification. In machine learning and deep learning, scene classification is typically regarded as a multi-classification problem, where the model needs to learn to map input features to predefined scene categories. Formula 3 can be used to calculate the probability that the input features belong to each scene category. By calculating the probability distribution, formula 3 can accurately classify the input features into the most likely scene category, improving the accuracy of scene classification.
[0106] Formula 4 selects the scene with the highest probability as the final predicted target scene by comparing the predicted probabilities of each scene category. This process ensures the certainty of the model output and avoids ambiguous or uncertain prediction results. By comparing the probabilities to determine the final scene, formula 4 actually implements a "majority voting" or "optimal selection" mechanism within the model, which helps reduce the impact of noise or outliers on the classification results and enhances the reliability of the classification results. The calculation process of formula 4 is relatively simple and efficient, and can determine the final scene in a short time, which is particularly important for application scenarios with high real-time requirements.
[0107] S103. The guide generates an updated prompt word sequence based on the predicted target scene and the original prompt word sequence.
[0108] In this step, the guide can generate an updated prompt word sequence based on the predicted target scene and the original prompt word sequence. The updated prompt word sequence includes the original prompt word sequence and an extended prompt word sequence. The extended prompt word sequence is a prompt word sequence that extends the original prompt word sequence in the direction of the predicted target scene.
[0109] Specifically, after determining the predicted target scene, the guide generates an updated prompt word sequence based on this scene and the original prompt word sequence. The updated prompt word sequence includes the original prompt word sequence and an extended prompt word sequence. The extended prompt word sequence is a set of words that extends the original prompt word sequence in the direction of the predicted target scene. These extended words are closely related to the specific scene and can enrich the vocabulary selection for content generation, improving the diversity and expandability of the content. The guide obtains the target scene extended word library from the preset database according to the predicted target scene. This word library contains common words and expressions related to specific scenes.
[0110] Then, the guide determines the similarity between each original prompt word in the original prompt word sequence and each extended word in the target scene extended word library. Similarity calculation can use methods such as cosine similarity and Jaccard similarity. The extended words with similarity greater than the preset similarity threshold are added to the extended prompt word sequence. Through this step, it is ensured that the extended prompt words are highly relevant to the original prompt information and closely centered around the target scene.
[0111] Optionally, a database can be preset in the system to store the extended thesaurus for different scenarios. Each scenario corresponds to an independent thesaurus, which contains common words and expressions related to that scenario. Among them, the construction of the thesaurus can be achieved through technical means such as manual collation, web crawler scraping, and natural language processing.
[0112] The booter adopts appropriate similarity calculation algorithms, such as cosine similarity, Jaccard similarity, etc., to evaluate the similarity between the original prompt words and the extended words in the extended thesaurus of the target scenario. Before calculating the similarity, the original prompt words and the extended words need to be converted into vector forms. This can be achieved through word embedding techniques (such as Word2Vec, GloVe, etc.), which map words into a high-dimensional vector space. Using the selected similarity algorithm, calculate the similarity scores between each original prompt word in the original prompt word sequence and each extended word in the extended thesaurus of the target scenario.
[0113] To control the quality of the extended words, the system presets a similarity threshold. Only when the similarity score between the extended word and the original prompt word exceeds this threshold, is it regarded as a valid extended word. The booter traverses each extended word in the extended thesaurus of the target scenario and filters out the extended words that meet the conditions according to the similarity scores. Add the filtered extended words to the extended prompt word sequence according to certain rules (such as sorting by similarity scores from high to low). The extended prompt word sequence now contains the original prompt word sequence and the newly added extended words, forming a more rich and diverse vocabulary set. The booter combines the original prompt word sequence and the extended prompt word sequence to generate the final updated prompt word sequence. This sequence not only retains the original information input by the user but also adds extended vocabulary closely related to a specific scenario.
[0114] After generating the updated prompt word sequence, the booter can further optimize it, such as removing duplicate words, adjusting the order of words, etc., to improve the quality and usability of the sequence. As new data is continuously added and user feedback accumulates, the system should regularly update the extended thesaurus of the target scenario to reflect the changes in language and the actual needs of users. In addition, a user feedback mechanism can be introduced to allow users to evaluate and modify the generated extended prompt word sequence, and this feedback information can be used to further optimize the thesaurus and the similarity calculation algorithm.
[0115] In a possible embodiment, the booter first performs word segmentation on the original prompt information to obtain the original prompt word sequence: "Write / intelligent grid / maintenance / of / technology / report".
[0116] Next, the booter inputs the original prompt word sequence into the preset scenario classification model. In this example, the scenario classification model identifies that the predicted target scenario is "intelligent grid maintenance".
[0117] According to the predicted target scenario "smart grid maintenance", the guide retrieves the relevant extended vocabulary from the preset database. This extended vocabulary contains professional words and technical terms closely related to smart grid maintenance, such as "power grid monitoring", "fault diagnosis", "automated inspection", "energy management", etc.
[0118] The guide then calculates the similarity between each word in the original prompt word sequence and each word in the extended vocabulary. In the power field, the calculation of similarity may need to consider the professionalism of the vocabulary and the context. For example, it is found that the similarity between "smart grid" and words such as "power grid monitoring" and "fault diagnosis" is relatively high because they all belong to the category of smart grid maintenance and management. While the similarity between "report" and these words is relatively low because it is a more general vocabulary. The guide sets a preset similarity threshold and adds the words with similarity greater than this threshold to the extended prompt word sequence. In this example, assume that the similarity of "power grid monitoring" and "fault diagnosis" exceeds the threshold, so they are added to the extended prompt word sequence.
[0119] Finally, the guide merges the original prompt word sequence and the extended prompt word sequence to generate an updated prompt word sequence: "Write a technical report on / smart grid / maintenance / power grid monitoring / fault diagnosis". This sequence not only retains the user's original intention (writing a technical report on smart grid maintenance), but also adds professional words and technical terms closely related to the smart grid maintenance scenario, providing more specific and professional materials for subsequent content generation.
[0120] S104. The guide inputs the original prompt information and the updated prompt word sequence into the artificial intelligence content generation engine to generate the artificial intelligence generated content corresponding to the prompt information.
[0121] In this step, the guide inputs the original prompt information and the updated prompt word sequence into the artificial intelligence content generation engine. Among them, this engine can be an existing commercial general artificial intelligence content generation engine or an artificial intelligence content generation engine developed separately for various scenarios.
[0122] In this embodiment, the original prompt information input by the user is obtained through a bootstrapper, and the original prompt information is segmented to generate an original prompt word sequence. Then, the original prompt word sequence is input into a preset scenario classification model to determine the predicted target scenario corresponding to the original prompt information. Thus, an updated prompt word sequence is generated based on the predicted target scenario and the original prompt word sequence. Next, the original prompt information and the updated prompt word sequence are input into an artificial intelligence content generation engine to generate artificial intelligence generated content corresponding to the prompt information. Thereby, the input original prompt information is subjected to targeted expansion in a scenario, so as to more accurately identify the user's intention and determine the target scenario that matches it, ensuring that the generated content better meets the requirements of a specific scenario, improving the pertinence and practicality of the content. Moreover, the requirement for the user's expression ability is reduced, the number of interaction rounds is reduced, and the generation efficiency and accuracy of the artificial intelligence generated content are effectively improved.
[0123] Figure 2 FIG. is a schematic flow chart of a scenario-based AIGC content generation method according to another exemplary embodiment of the present application. As Figure 2 shown, the scenario-based AIGC content generation method provided in this embodiment includes:
[0124] S201. Obtain the original prompt information input by the user through a bootstrapper, and segment the original prompt information to generate an original prompt word sequence.
[0125] In this step, the original prompt information input by the user is obtained through a bootstrapper. The bootstrapper segments the original prompt information, cuts the continuous text into independent lexical units, and generates an original prompt word sequence. Word segmentation is a basic task in natural language processing, which helps in subsequent understanding and processing of the prompt information.
[0126] Specifically, to provide a good user experience, the user can input the required original prompt information through the input interface of a terminal device (such as a computer, mobile phone, tablet, etc.). Among them, various input methods such as a text box and a voice input button can be provided on the interface to meet the needs of different users. For example, for users with poor eyesight, a voice input function can be provided; for users who prefer keyboard input, a clear text box is provided. The bootstrapper is responsible for receiving the original prompt information input by the user through the interface. This information will be temporarily stored in the local cache or memory for subsequent processing.
[0127] S202. The bootstrapper inputs the original prompt word sequence into a preset scenario classification model to determine the predicted target scenario corresponding to the original prompt information.
[0128] In this step, the booter inputs the original prompt sequence into a preset scenario classification model to determine the predicted target scenario corresponding to the original prompt information. The preset scenario classification model is a classifier generated based on a deep learning algorithm.
[0129] S203. The booter obtains the target scenario extension vocabulary from the preset database according to the predicted target scenario.
[0130] Specifically, the system needs to clarify the predicted target scenario of the content to be generated currently. This scenario may be determined based on various situations such as user input, context analysis, or system presets. The booter accesses a pre-constructed and maintained database according to the determined predicted target scenario. This database contains extension vocabularies associated with various possible scenarios, and each vocabulary is a collection of carefully selected words, phrases, or concepts for a specific scenario. Retrieve and extract the extension vocabulary directly related to the predicted target scenario from the database to provide a rich vocabulary resource for the subsequent steps.
[0131] S204. The booter determines the similarity between each original prompt word in the original prompt sequence and each extension word in the target scenario extension vocabulary.
[0132] In this step, the booter determines the similarity between each original prompt word in the original prompt sequence and each extension word in the target scenario extension vocabulary to add the extension words with similarity greater than the preset similarity threshold to the extended prompt sequence.
[0133] Specifically, the booter analyzes each word in the input original prompt sequence one by one to understand the meaning and context association of each prompt word. Then, using natural language processing techniques such as cosine similarity, Jaccard similarity coefficient, or more advanced semantic similarity algorithms, calculate the similarity between the original prompt word and each extension word in the target scenario extension vocabulary. Then, set a similarity threshold and select all the extension words with similarity greater than this threshold to form the extended prompt sequence. This step ensures that only the extension words highly relevant to the original prompt words and suitable for the target scenario are selected.
[0134] S205. The booter generates an updated prompt sequence according to the original prompt sequence and the extended prompt sequence.
[0135] Specifically, the booter intelligently integrates the selected extended prompt words with the original prompt sequence. This may include directly inserting, replacing some original words, or constructing new phrase structures around the original words to ensure that the updated prompt sequence not only retains the original intention but also enriches the scenario details.
[0136] Finally, the guide outputs a constructed and optimized updated prompt sequence as the input to the AI content generation engine, thereby triggering the model to generate content output that highly matches the target scenario and is targeted.
[0137] Furthermore, the guide can display a prompt network diagram through the terminal device. The prompt network diagram includes the original prompt node sequence corresponding to the original prompt sequence, the extended prompt node sequence corresponding to the extended prompt sequence, and the feature edge sequence for establishing the mapping relationship between each original prompt node in the original prompt node sequence and each extended prompt node in the extended prompt node sequence. By obtaining the weight adjustment instruction for the target feature edge in the feature edge sequence through the guide, only the extended prompt corresponding to the target feature edge with a weight greater than the first preset weight threshold after adjustment is determined as the target extended prompt. The updated prompt sequence is generated according to the original prompt sequence and all the target extended prompts.
[0138] Specifically, the guide first displays a prompt network diagram through the terminal device. This network diagram is a visual tool for showing the relationship between the original prompt sequence (original prompt node sequence) and the extended prompt sequence (extended prompt node sequence). Each prompt is represented as a node, and their relationship is represented by feature edges. The feature edges not only connect the original prompts and the extended prompts but also carry weight information, which reflects the closeness or importance of their relationship. Next, the user can interact with the prompt network diagram through the interface on the terminal device. Specifically, the user can select the feature edges in the network diagram and input the weight adjustment instructions for these feature edges through the guide. These instructions can be to increase, decrease, or keep the weight unchanged. The guide will update the weights of the feature edges in real time according to the user's instructions. During the process of adjusting the weights, a first preset weight threshold is set. Only when the weight of the feature edge is adjusted to exceed this threshold, the corresponding extended prompt will be regarded as the target extended prompt, that is, those prompts that have an important impact or are highly relevant to the generated content.
[0139] Finally, the guide will generate an updated prompt sequence according to the original prompt sequence and all the nodes selected as the target extended prompts. This updated sequence not only contains the user's initial intention (expressed through the original prompt sequence) but also incorporates the additional information and details emphasized by the user through adjusting the weights of the feature edges.
[0140] Furthermore, after obtaining the weight adjustment instruction for the target feature edge in the feature edge sequence through the guide, the guide can also add the original prompt corresponding to the target feature edge with a weight greater than the second preset weight threshold after adjustment to the target scenario extension word library.
[0141] In the above solution, by adding the original prompt words with higher weights in the user feedback to the target scenario extension word library, the system can optimize and expand the content of its word library in real time and dynamically. This process ensures that the word library can keep up with user needs and language development trends, improving the timeliness and relevance of the system. As the word library continues to be enriched and optimized, the system can more accurately understand and match the user's intentions and scenario requirements when generating AIGC content, which helps to improve the intelligent level of content generation and makes the generated content more in line with the user's expectations and requirements. In addition, by continuously absorbing user feedback to optimize itself, the system can promote the expansion and optimization of the word library. When facing diverse user needs and complex language environments, the system can exhibit higher robustness and adaptability.
[0142] It is worth noting that the second preset weight threshold can also be set to be greater than the first preset weight threshold. By setting two different levels of weight thresholds, the importance levels of the prompt words can be more finely divided. The first preset weight threshold is used to screen out the basic prompt words that have an important impact on content generation, while the second preset weight threshold further screens out those prompt words with higher relevance or importance.
[0143] By setting the second preset weight threshold to be higher than the first preset weight threshold, the system can screen out the original prompt words that are more closely related to the predicted target scenario. And add the original prompt words corresponding to the target feature edges that are adjusted to be greater than the second preset weight threshold to the target scenario extension word library, which will play a more core role in the subsequent content generation process, ensuring that the generated content is closer to the user's intentions and specific scenario requirements, thereby improving the pertinence and relevance of the content.
[0144] During the user interaction process, the system dynamically adds the original prompt words with higher weights to the target scenario extension word library according to the user's adjustment instructions for the feature edge weights. Since the second preset weight threshold is higher than the first preset weight threshold, only those prompt words that are truly valued by the user and highly relevant to the scenario can be included in the word library. This mechanism not only ensures the quality of the word library content but also promotes the continuous optimization and update of the word library, enabling it to better adapt to the changes in different scenarios and user needs.
[0145] It is worth noting that in a possible application scenario, the architecture of the scenario-based content generation system may include a cluster of terminal devices and a server. Among them, each terminal device in the cluster of terminal devices is communicatively connected to the server through a local area network, the server is communicatively connected to an artificial intelligence content generation engine set on a wide area network, a preset database is set in the server, and a bootstrap is set in the terminal device and / or the server.
[0146] Specifically, the terminal device cluster can be composed of multiple terminal devices, which can be computers, mobile phones, tablets, etc. Users input the original prompt information or perform other interaction operations through these devices. The server, as the core component of the system, is responsible for handling the communication between the terminal devices and the artificial intelligence content generation engine, and at the same time stores and manages the preset database. The server is connected to the terminal device cluster through a local area network to ensure the efficiency and security of data transmission. The artificial intelligence content generation engine is deployed on the wide area network and has powerful content generation capabilities. The engine receives the updated prompt word sequence from the server and generates high-quality content based on it. The preset database is used to store key data such as the scenario expansion word library, providing necessary scenario knowledge and vocabulary resources for the guide. The database is located in the server for easy access and update. In addition, the guide can be deployed on the terminal device or the server, responsible for generating the updated prompt word sequence according to the predicted target scenario and the original prompt word sequence, and passing the result to the artificial intelligence content generation engine.
[0147] It is worth noting that the architecture of the above-mentioned scenario-based content generation system can be applied to the content creation and confidentiality collaboration platform within an enterprise. Specifically, in some enterprise-level applications, especially in the field of content creation involving sensitive information or trade secrets, such as financial analysis, legal document writing, product R & D reports, etc., enterprises need to build an internal local area network to ensure the security and confidentiality of data. At the same time, in order to improve the efficiency and quality of content creation, enterprises also hope to use external artificial intelligence content generation technology to assist in creation. Then an internal local area network is deployed within the enterprise, and all terminal devices (such as employee computers, servers, etc.) communicate through the local area network to ensure the security and speed of data transfer within the enterprise. Strict access control and data encryption mechanisms are set up within the local area network to prevent external illegal access and data leakage. Enterprises have strict confidentiality requirements for the original data, intermediate processes, and final results of content creation. It is necessary to ensure that no sensitive information is leaked to the external network or third parties during the content creation process. However, enterprises hope to use external artificial intelligence content generation engines to assist in content creation to improve creation efficiency and quality. For example, using the artificial intelligence content generation engine to generate preliminary copywriting, report frameworks, or data analysis charts, etc., and then having internal enterprise personnel review and modify them.
[0148] Therefore, terminal devices such as computers of internal employees in the enterprise are all connected to the local area network, which are used for inputting original prompt information, viewing generated content, etc. Then, through the server, which is the core component within the local area network, it stores and manages the preset database, processes requests from terminal devices, and communicates with the external artificial intelligence content generation engine. Then it accesses the artificial intelligence content generation engine deployed on the external wide area network, but interacts through the secure communication interface with the internal server of the enterprise, receives the updated prompt word sequence from the server, and returns the generated content. A secure communication channel is established between the server and the external artificial intelligence content generation engine, such as using the SSL / TLS encryption protocol. The data transmitted is encrypted to ensure that it is not stolen or tampered with during the transmission process. Strict access control is carried out for the external artificial intelligence content generation engine, and only authorized internal servers can communicate with it.
[0149] Specifically, the above-mentioned enterprise-level internal content creation and confidentiality collaboration platform can be applied to the financial industry, government agencies, and the healthcare industry. Among them, for the financial industry, financial institutions such as banks and insurance companies need to process a large amount of sensitive information, such as customer information, transaction records, etc. Using this system architecture, while ensuring data confidentiality, it can efficiently generate customized financial reports, market analysis, product promotion, and other content. For government agencies, government departments need to process a large amount of internal documents and sensitive information, and at the same time need to release policy interpretations, announcements, and other content externally. This system architecture can help government departments efficiently generate and release customized content under the premise of confidentiality. For the healthcare industry, hospitals, research institutions, etc. need to process sensitive information such as patients' medical records and research results. Using this system architecture, while protecting patients' privacy and research results, it can efficiently generate medical reports, research papers, and other content.
[0150] Correspondingly, the above-mentioned booter inputs the original prompt information and the updated prompt word sequence into the preset artificial intelligence content generation engine. Specifically, the booter can input the original prompt information and the updated prompt word sequence into the artificial intelligence content generation engine through the wide area network. Correspondingly, after the artificial intelligence generated content corresponding to the generated prompt information is generated, the artificial intelligence content generation engine can also send the artificial intelligence generated content to the server for display on the terminal device.
[0151] Specifically, the user can input the original prompt information through the terminal device, and this information may include key elements such as the theme, style, and target audience of content generation. Then, the terminal device sends the original prompt information to the server. The bootstrapper in the server obtains the corresponding target scenario extension word library from the preset database according to the predicted target scenario. The bootstrapper calculates the similarity between the original prompt words and the words in the extension word library, and filters out the extension words that are highly relevant to the original prompt words and suitable for the target scenario. The bootstrapper integrates the filtered extension words with the original prompt words to generate an updated prompt word sequence.
[0152] The bootstrapper sends the updated prompt word sequence to the artificial intelligence content generation engine through the wide area network. This process relies on a stable network connection and an efficient data transmission protocol to ensure the accurate transmission of information. After receiving the updated prompt word sequence, the artificial intelligence content generation engine uses its powerful generation ability to output high-quality content. This content may include various forms such as text, images, and audio. The generated content is sent back to the server. The server distributes the content to each terminal device in the corresponding terminal device cluster. The terminal device receives and displays the generated content for the user to view and use.
[0153] S206. The bootstrapper inputs the original prompt information and the updated prompt word sequence into the artificial intelligence content generation engine to generate the artificial intelligence generated content corresponding to the prompt information.
[0154] In this step, the bootstrapper inputs the original prompt information and the updated prompt word sequence into the artificial intelligence content generation engine. Among them, the engine can be an existing commercial general artificial intelligence content generation engine or an artificial intelligence content generation engine developed separately for various scenarios.
[0155] Figure 3 is a schematic structural diagram of a scenario-based content generation system according to an exemplary embodiment of the present invention. As Figure 3 shown, the scenario-based content generation system 300 provided in this embodiment includes: a bootstrapper 310 and an artificial intelligence content generation engine 320;
[0156] Obtain the original prompt information input by the user through the bootstrapper 310, and perform word segmentation processing on the original prompt information to generate an original prompt word sequence;
[0157] The bootstrapper 310 inputs the original prompt word sequence into a preset scenario classification model to determine the predicted target scenario corresponding to the original prompt information. The preset scenario classification model is a classifier generated based on a deep learning algorithm;
[0158] The guide 310 generates an updated prompt word sequence according to the predicted target scenario and the original prompt word sequence. The updated prompt word sequence includes the original prompt word sequence and an extended prompt word sequence. The extended prompt word sequence is a prompt word sequence obtained by extending the original prompt word sequence in the direction of the predicted target scenario;
[0159] The guide 310 inputs the original prompt information and the updated prompt word sequence into the artificial intelligence content generation engine 320 to generate artificial intelligence generated content corresponding to the prompt information.
[0160] Optionally, the guide 310 generating an updated prompt word sequence according to the predicted target scenario and the original prompt word sequence includes:
[0161] The guide 310 obtains a target scenario extension word library from a preset database according to the predicted target scenario;
[0162] The guide 310 determines the similarity between each original prompt word in the original prompt word sequence and each extension word in the target scenario extension word library, and adds the extension words with a similarity greater than a preset similarity threshold to the extended prompt word sequence;
[0163] The guide 310 generates the updated prompt word sequence according to the original prompt word sequence and the extended prompt word sequence.
[0164] Optionally, the scenario-based content generation system includes a terminal device cluster and a server. Each terminal device in the terminal device cluster is communicatively connected to the server through a local area network. The server is communicatively connected to the artificial intelligence content generation engine 320 provided on a wide area network. The preset database is provided in the server, and the guide 310 is provided in the terminal device and / or the server;
[0165] Correspondingly, the guide 310 inputting the original prompt information and the updated prompt word sequence into the preset artificial intelligence content generation engine 320 includes:
[0166] The guide 310 inputs the original prompt information and the updated prompt word sequence into the artificial intelligence content generation engine 320 through the wide area network;
[0167] Correspondingly, after generating the artificial intelligence generated content corresponding to the prompt information, it further includes:
[0168] The artificial intelligence content generation engine 320 sends the artificial intelligence generated content to the server for display on the terminal device.
[0169] Optionally, the guide 310 generates the updated prompt sequence according to the original prompt sequence and the extended prompt sequence, including:
[0170] The guide 310 displays a prompt word network diagram through the terminal device, where the prompt word network diagram includes an original prompt word node sequence corresponding to the original prompt sequence, an extended prompt word node sequence corresponding to the extended prompt sequence, and a feature edge sequence for establishing a mapping relationship between each original prompt word node in the original prompt word node sequence and each extended prompt word node in the extended prompt word node sequence;
[0171] Obtain a weight adjustment instruction for a target feature edge in the feature edge sequence through the guide 310, and determine only the extended prompt words corresponding to the target feature edges with weights greater than a first preset weight threshold after adjustment as target extended prompt words;
[0172] Generate the updated prompt sequence according to the original prompt sequence and all the target extended prompt words.
[0173] Optionally, after obtaining the weight adjustment instruction for the target feature edge in the feature edge sequence through the guide 310, it further includes:
[0174] The guide 310 adds the original prompt words corresponding to the target feature edges with weights greater than a second preset weight threshold after adjustment to the target scenario extended word library.
[0175] Optionally, the preset scenario classification model includes a prompt word input layer, a convolutional neural network layer, a long short-term memory network layer, and a scenario output layer;
[0176] The prompt word input layer is used to input the original prompt sequence and convert the original prompt sequence into an original prompt word vector sequence to transfer the original prompt word vector sequence to the convolutional neural network layer;
[0177] The convolutional neural network layer is used to extract features from the original prompt word vector sequence to generate an original prompt word feature sequence and transfer the original prompt word feature sequence to the long short-term memory network layer;
[0178] The long short-term memory network layer is used to generate a hidden state sequence according to the original prompt word feature sequence and transfer the hidden state sequence to the scenario output layer;
[0179] The scenario output layer inputs the hidden state sequence into a preset classifier to output a predicted candidate scenario sequence, where the predicted candidate scenario sequence includes at least one candidate scenario sequence, and the predicted candidate scenario sequence includes the predicted target scenario.
[0180] Optionally, the convolutional neural network layer is used to extract features from the original prompt word vector sequence to generate an original prompt word feature sequence, including:
[0181] The convolutional neural network layer uses Formula 1 and generates the original prompt word feature sequence Y according to the original prompt word vector sequence X, where Formula 1 is:
[0182] Y = σ(W * X + b)
[0183] where σ is an activation function, W is the weight matrix of the convolution kernel, * is the convolution operation, and b is the bias term;
[0184] Correspondingly, the long short-term memory network layer is used to generate a hidden state sequence according to the original prompt word feature sequence, including:
[0185] The long short-term memory network layer uses Formula 2 and generates the hidden state h at the current time step in the hidden state sequence H according to the original prompt word feature sequence Y t , where Formula 2 is:
[0186] h t = LSTM(h t-1 , Y)
[0187] where LSTM is the operation of the long short-term memory network layer, and h t-1 is the hidden state at the previous time step in the hidden state sequence H.
[0188] Figure 4 is a schematic structural diagram of an electronic device shown according to an exemplary embodiment of the present invention. As Figure 4 shown, an electronic device 400 provided in this embodiment includes: a processor 401 and a memory 402; wherein:
[0189] The memory 402 is used to store a computer program, and this memory can also be flash (flash memory).
[0190] The processor 401 is used to execute the execution instructions stored in the memory to implement each step in the above method. For specific reference, please refer to the relevant descriptions in the previous method embodiments.
[0191] Optionally, the memory 402 can be either independent or integrated with the processor 401.
[0192] When the memory 402 is a device independent of the processor 401, the electronic device 400 may further include:
[0193] A bus 403 for connecting the memory 402 and the processor 401.
[0194] This embodiment also provides a readable storage medium storing a computer program, and when at least one processor of an electronic device executes the computer program, the electronic device executes the methods provided by the above various embodiments.
[0195] This embodiment also provides a program product, which includes a computer program stored in a readable storage medium. At least one processor of the electronic device can read the computer program from the readable storage medium, and the execution of the computer program by at least one processor enables the electronic device to implement the methods provided by the above various embodiments.
[0196] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present application are pointed out by the claims.
[0197] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A scenario-based AIGC content generation method, characterized in that, Complete AIGC content generation using a scenario-based content generation system, where the scenario-based content generation system includes a guide and an artificial intelligence content generation engine; Specifically, it includes the following steps: S101: Obtain the original prompt information input by the user through the guide, perform word segmentation on the original prompt information, and generate an original prompt word sequence; S102: The guide inputs the original prompt word sequence into a preset scenario classification model to determine the predicted target scenario corresponding to the original prompt information. The preset scenario classification model is a classifier generated based on a deep learning algorithm; S103: The guide generates an updated prompt word sequence according to the predicted target scenario and the original prompt word sequence. The updated prompt word sequence includes an original prompt word sequence and an extended prompt word sequence. The extended prompt word sequence is a prompt word sequence obtained by expanding the original prompt word sequence in the direction of the predicted target scenario; S104: The guide inputs the original prompt information and the updated prompt word sequence into the artificial intelligence content generation engine to generate artificial intelligence-generated content corresponding to the prompt information.
2. The method for generating scenario-based AIGC content according to claim 1, wherein In step S103, the guide generates an updated prompt word sequence according to the predicted target scenario and the original prompt word sequence, including: The guide obtains a target scenario extension word library from a preset database according to the predicted target scenario; The guide determines the similarity between each original prompt word in the original prompt word sequence and each extension word in the target scenario extension word library, and adds the extension words with a similarity greater than a preset similarity threshold to the extended prompt word sequence; The guide generates the updated prompt word sequence according to the original prompt word sequence and the extended prompt word sequence.
3. The method for generating scene-based AIGC content according to claim 2, wherein, The guide generates the updated prompt word sequence according to the original prompt word sequence and the extended prompt word sequence, including: The guide displays a prompt word network diagram through the terminal device, where the prompt word network diagram includes an original prompt word node sequence corresponding to the original prompt word sequence, an extended prompt word node sequence corresponding to the extended prompt word sequence, and a feature edge sequence for establishing a mapping relationship between each original prompt word node in the original prompt word node sequence and each extended prompt word node in the extended prompt word node sequence; Obtain a weight adjustment instruction for a target feature edge in the feature edge sequence through the guide, and only determine the extended prompt word corresponding to the target feature edge with a weight greater than a first preset weight threshold after adjustment as the target extended prompt word; Generate the updated prompt word sequence according to the original prompt word sequence and all the target extended prompt words.
4. The method for generating scenario-based AIGC content according to claim 3, wherein After obtaining the weight adjustment instruction for the target feature edge in the feature edge sequence through the guide, it further includes: The guide adds the original prompt word corresponding to the target feature edge with a weight greater than a second preset weight threshold after adjustment to the target scenario extension word library.
5. The method for generating scenario-based AIGC content according to claim 1, wherein The described scenario-based content generation system includes a cluster of terminal devices and a server. Each terminal device in the cluster of terminal devices is communicatively connected to the server via a local area network. The server is communicatively connected to the artificial intelligence content generation engine provided on a wide area network. The preset database is provided in the server, and the booter is provided in the terminal device and / or the server; The booter inputs the original prompt information and the updated prompt word sequence into the preset artificial intelligence content generation engine, including: The booter inputs the original prompt information and the updated prompt word sequence into the artificial intelligence content generation engine via the wide area network; After generating the artificial intelligence-generated content corresponding to the prompt information, the artificial intelligence content generation engine sends the artificial intelligence-generated content to the server for display on the terminal device.
6. The method for generating scene-based AIGC content according to claim 1, wherein, The preset scenario classification model includes a prompt word input layer, a convolutional neural network layer, a long short-term memory network layer, and a scenario output layer; The prompt word input layer is used to input the original prompt word sequence and convert the original prompt word sequence into an original prompt word vector sequence to transfer the original prompt word vector sequence to the convolutional neural network layer; The convolutional neural network layer is used to extract features from the original prompt word vector sequence to generate an original prompt word feature sequence and transfer the original prompt word feature sequence to the long short-term memory network layer; The long short-term memory network layer is used to generate a hidden state sequence based on the original prompt word feature sequence and transfer the hidden state sequence to the scenario output layer; The scenario output layer inputs the hidden state sequence into a preset classifier to output a predicted candidate scenario sequence. The predicted candidate scenario sequence includes at least one candidate scenario sequence, and the predicted candidate scenario sequence includes the predicted target scenario.
7. The method for generating scenario-based AIGC content according to claim 6, wherein The convolutional neural network layer is used to extract features from the original prompt word vector sequence to generate an original prompt word feature sequence, including: The convolutional neural network layer uses the following formula and generates the original prompt word feature sequence Y based on the original prompt word vector sequence X, where the formula is: Y = σ(W * X + b) where σ is the activation function, W is the weight matrix of the convolutional kernel, * is the convolution operation, and b is the bias term; The long short-term memory network layer is used to generate a hidden state sequence based on the original prompt word feature sequence, including: The long short-term memory network layer uses the following formula and generates the hidden state h at the current time step in the hidden state sequence H according to the original prompt word feature sequence Y t , where the formula is: h t = LSTM(h t-1 , Y) where LSTM is the operation of the long short-term memory network layer, and h t-1 is the hidden state at the previous time step in the hidden state sequence H.
8. A scenario-based AIGC content generation system, characterized in that, including: A booter and an artificial intelligence content generation engine; The booter obtains the original prompt information input by the user, performs word segmentation on the original prompt information to generate an original prompt word sequence; The booter inputs the original prompt word sequence into the preset scenario classification model to determine the predicted target scenario corresponding to the original prompt information. The preset scenario classification model is a classifier generated based on a deep learning algorithm; The booter generates an updated prompt word sequence according to the predicted target scenario and the original prompt word sequence. The updated prompt word sequence includes the original prompt word sequence and an extended prompt word sequence. The extended prompt word sequence is a prompt word sequence obtained by extending the original prompt word sequence in the direction of the predicted target scenario; The booter inputs the original prompt information and the updated prompt word sequence into the artificial intelligence content generation engine to generate the artificial intelligence generated content corresponding to the prompt information.