AIGC-based Content Creation and Multimodal Display Platform
Through the content creation and multimodal display platform based on AIGC, the interaction management and user behavior analysis modules are used to generate standard intention data and retrieval enhancement data sets, which improves the accuracy of intention recognition and content generation, solves the shortcomings of intention understanding and personalized recommendation in the existing technology, and achieves efficient personalized content generation.
Patent Information
- Application Number
- CN202411688064.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-25
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-11-25
AI Technical Summary
The prior art has shortcomings in intention recognition and personalized recommendations, especially in complex intention understanding, which is difficult to flexibly adjust, resulting in limited system understanding capabilities and the accuracy of personalized content generation is also limited by computing resources, affecting the quality of recommendations and user experience.
A content creation and multimodal display platform based on AIGC is adopted, and standard instruction word vectors and context information are obtained through the interaction management module to generate standard intention data, and a search enhancement data set is obtained in line with the user behavior analysis module. The Transformer architecture and self-attention mechanism are used to improve the accuracy of intention recognition, and multimodal content data is generated through the content creation module.
It improves the accuracy of user intention interpretation, can generate highly relevant multimodal content data for individual users, reduces the computing burden of content creation, and provides better user interaction experience and personalized content generation.
Smart Images

Figure CN119538924B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a content creation and multi-modal display platform based on AIGC. Background Art
[0002] With the explosive growth of digital information, users have put forward higher requirements for diversified and personalized content creation methods. Traditional methods rely on manual editing and the generation of content in a single modality (such as pure text or pictures). The content generation process is time-consuming and laborious, and there are problems of low efficiency and strong dependence on professional experience.
[0003] To address these challenges, in recent years, researchers and enterprises have begun to explore content generation technologies based on AIGC (Generative Artificial Intelligence), extracting useful information from massive data through machine learning and natural language processing technologies, and automatically generating high-quality, multi-modal content to meet the diverse needs of users.
[0004] However, there are still some deficiencies in the existing technologies in terms of intent recognition and personalized recommendation. Especially in the understanding of complex intents, existing systems often rely on fixed user input parsing patterns and are difficult to adjust flexibly, resulting in limited understanding capabilities of the systems. Moreover, the accuracy of personalized content generation is also restricted by computing resources. The huge amount of retrieved information increases the computational complexity, leading to slow speed in real-time response and generating personalized content, thus affecting the recommendation quality and user experience. Summary of the Invention
[0005] The present invention provides a content creation and multi-modal display platform based on AIGC to solve at least one of the problems mentioned in the above background art.
[0006] The specific technical solutions provided by this application are as follows:
[0007] A content creation and multi-modal display platform based on AIGC, comprising an interaction management module, a user behavior analysis module, and a content creation module that are communicatively connected;
[0008] The interaction management module is used to obtain a standard instruction word vector and a first retrieval enhanced data set according to the current input information, and generate standard intent data according to the standard instruction word vector and context information;
[0009] The user behavior analysis module is used to obtain a second retrieval enhanced data set according to the historical content description word vector of the user;
[0010] The content creation module is used to generate multi-modal content data and its corresponding content description word vector according to the standard intent data, the first retrieval enhanced data set, and the second retrieval enhanced data set;
[0011] The interactive management module includes an input processing module, a context management module, and an intent recognition module; the input processing module is used to process the current input information into a standard instruction word vector and a first retrieval enhancement dataset; the context management module is used to generate context information according to the dialogue history instruction word vector and the content description word vector; the intent recognition module is used to generate standard intent data according to the standard instruction word vector and the context information.
[0012] As a preferred solution, the input processing module includes an instruction input module and a retrieval enhancement input module; the instruction input module is used to obtain an original operation instruction according to the current input information and process the original operation instruction into a standard instruction word vector; the retrieval enhancement input module is used to obtain a first retrieval enhancement dataset according to the current input information, construct retrieval keywords corresponding to the first retrieval enhancement dataset, and store them.
[0013] As a preferred solution, the original operation instruction is a text instruction or a voice instruction; the instruction input module includes a voice-to-text conversion module and a first word vector extraction module; the voice-to-text conversion module is used to convert the voice instruction into a text instruction through a modality conversion strategy when the original operation instruction is a voice instruction; the first word vector extraction module is used to process the text instruction through a word vector extraction strategy to obtain the standard instruction word vector of the original operation instruction.
[0014] As a preferred solution, the modality conversion strategy specifically includes:
[0015] Dividing the voice instruction into several audio frames, and preprocessing each audio frame to obtain corresponding spectrum data;
[0016] Obtaining Mel spectrum data by passing the spectrum data through a Mel filter, performing cepstrum analysis based on the Mel spectrum data to obtain Mel frequency cepstral coefficients, and integrating the Mel frequency cepstral coefficients into the current audio feature;
[0017] Converting the current audio feature into a text instruction through a speech recognition model.
[0018] As a preferred solution, the Mel spectrum data is expressed as:
[0019] ,
[0020] ,
[0021] where represents the m-th Mel spectrum data, m is the number of the Mel filter; k represents the number of the frequency point, and N represents the total number of frequency points; represents the power spectrum of the k-th frequency point of the frequency domain signal; represents the value of the m-th Mel filter at the frequency point k; represents the central position of the m-th Mel filter on the linear spectrum;
[0022] The Mel-frequency cepstral coefficients are expressed as:
[0023] ,
[0024] where M represents the number of Mel filters; n represents the number of the Mel-frequency cepstral coefficients; represents the value of the n-th Mel-frequency cepstral coefficient.
[0025] As a preferred solution, the word vector extraction strategy specifically includes:
[0026] Using a word segmentation algorithm to process the text instruction to obtain several instruction word segments;
[0027] Inputting the several instruction word segments after preprocessing into the Skip-gram model to obtain the standard instruction word vector of the original operation instruction.
[0028] As a preferred solution, generating the standard intent data according to the standard instruction word vector and context information includes the steps of:
[0029] Integrating the standard instruction word vector and context information into several continuous word vector sequences and word vector type sequences according to the data timestamp; the word vector sequence consists of a start identifier, an end identifier, and several word segmentation IDs;
[0030] Inputting the word vector sequence and word vector type sequence into several layers of Transformer to generate the context representation of each word segmentation ID; extracting the start identifier from the last layer of Transformer as the comprehensive feature of the word vector sequence;
[0031] Outputting the intent recognition result with the highest probability through the fully connected layer and Softmax layer using the comprehensive feature as the standard intent data.
[0032] As a preferred solution, the user behavior analysis module includes an intent data management module, a portrait analysis module, and a data matching module; the intent data management module is used to store the standard intent data and generate the historical intent data of the user according to the standard intent data; the portrait analysis module is used to generate user preference features according to the historical content description word vector of the user; the data matching module is used to store the retrieved enhanced data and its corresponding feature index, and obtain the second retrieved enhanced data set by calculating the matching degree between the user preference features and the feature index.
[0033] As a preferred solution, the image analysis module stores a set of preference word vectors, and the set of preference word vectors stores a number of preference word vectors; generating user preference features based on the historical content description word vectors of the user specifically means generating user preference features based on the set of preference word vectors and the historical content description word vectors of the user; the user preference features store several groups of preference word vectors and their corresponding frequencies.
[0034] As a preferred solution, the content creation module includes a feature extraction module, a feature screening module, and a content generation module; the feature extraction module is used to obtain the data features of the first retrieval enhanced data set and the second retrieval enhanced data set; the feature screening module is used to remove the parts of the data features of the first retrieval enhanced data set and the second retrieval enhanced data set that are not relevant to the standard intent data to obtain an intent-related feature set; the content generation module encodes the standard intent data and the intent-related feature set into the same vector space through a Transformer encoder, and generates content data of the corresponding modality and its corresponding content description word vectors through a number of integrated generation models.
[0035] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0036] In this application, the interaction management module obtains the standard instruction word vector according to the current input information, and generates the standard intent data according to the standard instruction word vector and the context information, making it easier for the system to understand different user instructions and improving the accuracy of user intent interpretation; the user behavior analysis module obtains the second retrieval enhanced data set that meets the user's preferences according to the historical content description word vectors of the user; the content creation module generates multi-modal content data and its corresponding content description word vectors according to the standard intent data, the first retrieval enhanced data set input currently, and the second retrieval enhanced data set that meets the user's preferences. This application improves the accuracy of user intent interpretation, can generate highly relevant multi-modal content data for individual users to provide personalized content generation, and reduces the computational burden of content creation.
[0037] This application captures the complex dependencies between words through the Transformer architecture and the self-attention mechanism, enhancing the understanding depth and prediction accuracy of the dialogue system in intent recognition; through the fully connected layer for feature combination and dimensionality reduction and through the Softmax layer to output the intent recognition result with the highest probability as the standard intent data, enabling the system to effectively process complex semantic tasks and ultimately providing more accurate intent parsing and a better user interaction experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The drawings here are incorporated into the specification and form a part of this specification, indicating the embodiments that conform to the present invention, and are used together with the specification to explain the principles of the present invention.
[0039] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for description in the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0040] Figure 1 Structural schematic diagram of a content creation and multi-modal display platform based on AIGC provided by an embodiment of the present invention;
[0041] Figure 2 Structural schematic diagram of an interaction management module provided by an embodiment of the present invention;
[0042] Figure 3 Flow schematic diagram for generating standard intent data according to standard instruction word vectors and context information provided by an embodiment of the present invention;
[0043] Figure 4 Structural schematic diagram of a user behavior analysis module provided by an embodiment of the present invention;
[0044] Figure 5 Structural schematic diagram of a content creation module provided by an embodiment of the present invention. Detailed implementation manners
[0045] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0046] It should be noted that all directional indications (such as up, down, left, right, front, back...) in the embodiments of the present invention are only used to explain the relative position relationship and movement conditions between components in a specific posture (as shown in the drawings). If the specific posture changes, the directional indications will also change accordingly.
[0047] In addition, the descriptions involving "first", "second", etc. in the present invention are only for descriptive purposes and should not be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least the feature. In addition, the technical solutions between various embodiments may be combined with each other, but it must be based on the realization by those of ordinary skill in the art. When the combination of technical solutions is contradictory or unable to be realized, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0048] With the explosive growth of digital information, users have put forward higher requirements for diversified and personalized content creation methods. Traditional methods rely on manual editing and the generation of content in a single modality (such as pure text or pictures). The content generation process is time-consuming and laborious, with problems of low efficiency and strong dependence on professional experience.
[0049] To address these challenges, in recent years, researchers and enterprises have begun to explore content generation technologies based on AIGC (Generative Artificial Intelligence), extracting useful information from massive data through machine learning and natural language processing technologies, and automatically generating high-quality, multi-modal content to meet the diversified needs of users.
[0050] However, there are still some deficiencies in the existing technologies in intention recognition and personalized recommendation. Especially in the understanding of complex intentions, existing systems often rely on fixed user input parsing modes and are difficult to adjust flexibly, resulting in limited understanding capabilities of the systems. Moreover, the accuracy of personalized content generation is also restricted by computing resources. The huge amount of retrieved information increases the computational complexity, leading to slow speeds in real-time response and generating personalized content by the system, thus affecting the recommendation quality and user experience.
[0051] Therefore, the present invention provides a content creation and multi-modal display platform based on AIGC. The specific implementation process of the present invention will be elaborated in detail below.
[0052] Please refer to Figure 1 , the present invention provides a content creation and multi-modal display platform based on AIGC, including an interaction management module, a user behavior analysis module, and a content creation module that are communicatively connected;
[0053] The interaction management module is used to obtain a standard instruction word vector and a first retrieval enhancement data set according to the current input information, and generate standard intention data according to the standard instruction word vector and context information;
[0054] The user behavior analysis module is used to obtain a second retrieval enhancement data set according to the historical content description word vector of the user;
[0055] The content creation module is used to generate multi-modal content data and its corresponding content description word vectors based on standard intention data, a first retrieval augmentation dataset, and a second retrieval augmentation dataset;
[0056] Among them, the multi-modal content data can be text data, image data, audio data, and video data. Both the first retrieval augmentation dataset and the second retrieval augmentation dataset are used by the content creation module to achieve Retrieval-Augmented Generation. Retrieval-Augmented Generation is a model architecture that combines information retrieval and generation, used to improve the performance and accuracy of generative language models. Compared with relying solely on large-scale retrieval generators, this application reduces the computational burden of the content creation module by using the first retrieval augmentation dataset and the second retrieval augmentation dataset, and provides more personalized multi-modal content data for users.
[0057] In this application, the interaction management module obtains standard instruction word vectors according to the current input information, and generates standard intention data according to the standard instruction word vectors and context information, making it easier for the content creation module to understand different user instructions and improving the accuracy of user intention interpretation; the user behavior analysis module obtains a second retrieval augmentation dataset that conforms to the user's preferences according to the user's historical content description word vectors; the content creation module generates multi-modal content data and its corresponding content description word vectors according to the standard intention data, the currently input first retrieval augmentation dataset, and the second retrieval augmentation dataset that conforms to the user's preferences, and can generate highly relevant multi-modal content data for individual users to provide personalized content generation and reduce the computational burden of content creation.
[0058] Further, please refer to Figure 2 , the interaction management module includes an input processing module, a context management module, and an intention recognition module. The input processing module is used to process the current input information into standard instruction word vectors and a first retrieval augmentation dataset; the context management module is used to generate context information according to the dialogue history instruction word vectors and content description word vectors; the intention recognition module is used to generate standard intention data according to the standard instruction word vectors and context information.
[0059] The dialogue history instruction word vectors are all the standard instruction word vectors generated by the input processing module in this round of dialogue. Among them, this round of dialogue refers to the continuous question-and-answer or instruction exchange process between the user and the system in the same interaction. This round of dialogue starts from the user's first input of an instruction until the user ends the current session or no longer continues to input. In this round of dialogue, all the standard instruction word vectors generated by the input processing module constitute the dialogue history instruction word vectors.
[0060] In the interaction management module of this embodiment, the input processing module is responsible for receiving and processing the information currently input by the user, and processing the information input by the user (such as text, voice, etc.) into a standard instruction word vector and a first retrieval enhancement dataset. The standard instruction word vector is an identifier of the specific intention of the user, while the first retrieval enhancement dataset provides rich relevant background information. The context management module is responsible for managing the context of the conversation, summarizing all the historical word vector instructions and the generated multimodal content data in the current conversation turn to form complete context information, providing a data basis for the system to understand the evolution of the user's intention in a long-term conversation. The intention recognition module is responsible for generating intention data. Based on the standard instruction word vector provided by the input processing module and the context information, it uses a set algorithm or model to generate standard intention data to accurately represent the specific needs and intentions of the current user.
[0061] Further, the input processing module includes an instruction input module and a retrieval enhancement input module; the instruction input module is used to obtain the original operation instruction according to the current input information and process the original operation instruction into a standard instruction word vector; the retrieval enhancement input module is used to obtain the first retrieval enhancement dataset according to the current input information, construct the retrieval keywords corresponding to the first retrieval enhancement dataset and store them.
[0062] Among them, the instruction input module obtains the original operation instruction in the current input information based on a preset instruction interface, and then identifies the original operation instruction input by the user and outputs a standard instruction word vector; the output result is a standard instruction word vector to ensure that subsequent processing modules can uniformly understand and use this information. The retrieval enhancement input module obtains the first retrieval enhancement dataset in the current input information based on a preset file interface, constructs the retrieval keywords corresponding to the first retrieval enhancement dataset according to the modality, file name, and content keywords of the first retrieval enhancement dataset, and then transmits the first retrieval enhancement dataset and its corresponding retrieval keywords to a specified storage system or cache for quick access. The first retrieval enhancement dataset can be data in modalities such as text, voice, image, and video.
[0063] In one embodiment, the original operation instruction is a text instruction or a voice instruction; the instruction input module includes a voice-to-text conversion module and a first word vector extraction module. The voice-to-text conversion module is used to convert the voice instruction into a text instruction through a modality conversion strategy when the original operation instruction is a voice instruction; the first word vector extraction module is used to process the text instruction through a word vector extraction strategy to obtain the standard instruction word vector of the original operation instruction.
[0064] Further, the modality conversion strategy specifically includes:
[0065] A1. Divide the voice instruction into several audio frames, and preprocess each audio frame to obtain the corresponding spectral data;
[0066] Divide the voice command into several audio frames. Specifically, divide the continuous voice command into several short time periods, and each period is recorded as an audio frame. In one embodiment, the length of each frame is 25 milliseconds. Preprocess each audio frame, including denoising, normalization, and fast Fourier transform operations. Denoising and normalization are used to improve the signal quality, reduce environmental noise and other interferences. The fast Fourier transform converts the time-domain signal of each audio frame into a frequency-domain signal to obtain the spectral data corresponding to each audio frame.
[0067] A2. Obtain Mel spectral data through Mel filters for the spectral data, perform cepstral analysis based on the Mel spectral data to obtain Mel frequency cepstral coefficients, and integrate the Mel frequency cepstral coefficients into the current audio feature;
[0068] Pass the spectral data through a set number of Mel filters, and the output is Mel spectral data. The Mel spectral data represents the energy distribution of the spectrum on the Mel scale. Then, perform a discrete cosine transform (DCT) on the Mel spectrum to obtain Mel frequency cepstral coefficients (MFCCs). These coefficients represent the characteristics of the voice signal, and the correlation is removed, which is convenient for subsequent processing. Finally, integrate the Mel frequency cepstral coefficients corresponding to all audio frames into the current audio feature. The current audio feature is represented in the form of a vector, and the current audio feature contains the main information that can describe the voice command.
[0069] Among them, the Mel spectral data is expressed as:
[0070] ,
[0071] ,
[0072] Among them, represents the m-th Mel spectral data, where m is the number of the Mel filter; k represents the number of the frequency point, and N represents the total number of frequency points; represents the power spectrum of the k-th frequency point of the frequency-domain signal; represents the value of the m-th Mel filter at the frequency point k; represents the central position of the m-th Mel filter on the linear spectrum.
[0073] The Mel frequency cepstral coefficients are expressed as:
[0074] ,
[0075] Among them, M represents the number of Mel filters; n represents the number of the Mel frequency cepstral coefficients; represents the value of the n-th Mel frequency cepstral coefficient;
[0076] A3. Convert the current audio feature into a text instruction through a speech recognition model;
[0077] Use the speech recognition model (such as a deep neural network) deployed in the instruction input module to process the audio feature. Among them, the speech recognition model has been trained with a large amount of speech and text data and can recognize and convert different phonemes. Input the current audio feature into the speech recognition model, and the model outputs the corresponding text sequence, and integrate the text sequence into a text instruction.
[0078] Furthermore, the word vector extraction strategy specifically includes:
[0079] B1. Use a word segmentation algorithm to process the text instruction to obtain several instruction word segments; among them, the text instruction can be the original operation instruction obtained according to the current input information, or the text instruction obtained by the speech-to-text conversion module converting the speech instruction through the modality conversion strategy. The word segmentation algorithm can be based on existing word segmentation methods such as Jieba and NLTK to perform word segmentation on the text instruction to obtain instruction word segments. Among them, after using the word segmentation algorithm to process the text instruction, operations such as removing punctuation and stop words can also be included.
[0080] B2. Input the representation of the several instruction word segments after preprocessing into the Skip-gram model to obtain the standard instruction word vector of the original operation instruction.
[0081] Among them, the Skip-gram model is an implementation method of the lightweight neural network model Word2Vec, which is used to predict the context word vector based on a given word and is suitable for generating word vector representations. The Skip-gram model includes an input layer, a hidden layer, and an output layer. Among them, the size of the input layer is set to a one-hot vector of the vocabulary length, the dimension of the hidden layer is a dense vector of embedding_size for word vector representation, and the size of the output layer is a softmax layer of the vocabulary length. Using the Skip-gram model in the input processing stage can effectively generate word vectors that can represent the semantics of words, thereby improving the semantic understanding and similarity calculation of words.
[0082] Please refer to Figure 3 , in one embodiment, the intention recognition module, generating the standard intention data according to the standard instruction word vector and the context information, includes the steps:
[0083] C1. Integrate the standard instruction word vectors and context information into several consecutive word vector sequences and word vector type sequences according to the data timestamp; the word vector sequences consist of a start identifier, an end identifier, and several token IDs; based on the foregoing, the context information is generated from the dialogue history instruction word vectors and content description word vectors, so the standard instruction word vectors and context information can be integrated into a unified format of word vector sequences in this step. Among them, the word vector sequences can correspond to a standard instruction word vector, a dialogue history instruction word vector, or a content description word vector. Specifically, the word vector sequences consist of a start identifier, an end identifier, and several token IDs. For example, the text content corresponding to the word vector sequence is "Generate a dynamic graph about climate change based on the text generated in the previous item, and the output format needs to be GIF", and the word vector sequence is represented as [101, 3306, 736,..., 102], where 101 corresponds to the start identifier, 102 corresponds to the end identifier, and the numbers in the word vector sequence are the preset token IDs. Based on the foregoing text content, the word vector type sequence is used to mark the sentence types to which the tokens belong. For example, the user input is recorded as type 0, and the system output is recorded as type 1.
[0084] In this step, by integrating the standard instruction word vectors and context information into a unified format of word vector sequences and type sequences, the module can process input data from different sources, ensure that the information is processed under the same structure, simplify subsequent processing, enhance the model's understanding of information types, and facilitate improving the accuracy of intent recognition.
[0085] C2. Input the word vector sequences and word vector type sequences into several layers of Transformer to generate the context representation of each token ID; extract the start identifier from the last layer of Transformer as the comprehensive feature of the word vector sequence.
[0086] In an exemplary embodiment, input the word vector sequences and word vector type sequences into the BERT model. The several Transformer layers in the BERT model capture the relationships between each word and other words in the input word vector sequence through the self-attention mechanism, enabling the model to focus on the context information in the input sequence. Through several Transformer layers, each token ID finally generates a context representation, that is, the embedding vector of this word, which contains the context information of the entire sentence or paragraph. Since in the BERT model, the context representation of the start identifier is designed to contain the semantic information of the entire sequence, the context representation of the start identifier (denoted as [CLS] or 101) is extracted from the output of the last layer of Transformer as the comprehensive feature of the entire input sequence or sentence.
[0087] C3. Output the intention recognition result with the highest probability of the comprehensive feature through the fully connected layer and the Softmax layer as the standard intention data.
[0088] In this step, the fully connected layer further processes and reduces the dimension of the comprehensive feature through linear transformation and non-linear activation functions. The Softmax layer is used to normalize the output of the fully connected layer, convert the result into the probability distribution of several intention recognition results, and then take the intention recognition result with the highest probability as the standard intention data.
[0089] In this embodiment, the Transformer architecture and the self-attention mechanism are used to capture the complex dependencies between words, improving the understanding depth and prediction accuracy of the dialogue system in intention recognition; through the fully connected layer for feature combination and dimensionality reduction and through the Softmax layer to output the intention recognition result with the highest probability as the standard intention data, enabling the system to effectively process complex semantic tasks and ultimately providing more accurate intention parsing and better user interaction experience.
[0090] Furthermore, please refer to Figure 4 , the user behavior analysis module includes an intention data management module, a portrait analysis module, and a data matching module; the intention data management module is used to store the standard intention data and generate the historical intention data of the user according to the standard intention data; the portrait analysis module is used to generate user preference features according to the historical content description word vectors of the user; the data matching module is used to store the retrieved enhanced data and its corresponding feature indexes, and obtain the second retrieved enhanced data set by calculating the matching degree between the user preference features and the feature indexes.
[0091] In this embodiment, the intention data management module stores the standard intention data generated by the interaction management module, creates and maintains the historical intention data corresponding to each user based on the standard intention data, to provide reliable data storage and management support, ensuring that all user intention data is effectively recorded and updated. The portrait analysis module stores the historical content description word vectors of the user, that is, the content description word vectors corresponding to the multi-modal content data provided by the content creation module for the user in all interaction data generated by the user and this platform, extracts user preference features according to the historical content description word vectors, and the user preference features can be used as the user portrait to enable the platform to better personalize content recommendations. The data matching module facilitates quick query and matching by storing and managing the retrieved enhanced data and its feature indexes; and calculates the matching degree between the user preference features and the feature indexes of the retrieved data, and screens out the most matching second retrieved enhanced data set.
[0092] In one embodiment, the image analysis module stores a set of preference word vectors, and the set of preference word vectors stores a number of preference word vectors; generating user preference features based on the historical content description word vectors of the user specifically means generating user preference features based on the set of preference word vectors and the historical content description word vectors of the user; the user preference features store several groups of preference word vectors and their corresponding frequencies.
[0093] Among them, the set of preference word vectors is stored in the database of the image analysis module. Each preference word vector in the set of preference word vectors represents a potential user interest or preference. The word vectors may be extracted from the content description word vectors output by the content creation module for all user groups through a set statistical algorithm, and can capture typical user tendencies or preferences. The user preference features are recorded in the form of preference word vectors and their corresponding frequencies to help the system understand the number of times a certain preference is selected or reflected by the user's historical behavior, so as to measure the significance of these preferences and help make more accurate decisions in personalized content recommendation.
[0094] In one embodiment, obtaining the second retrieval enhancement dataset by calculating the matching degree between the user preference features and the feature index includes the steps of:
[0095] Processing the user preference features and the feature index by the padding method; in this embodiment, both the user preference features and the feature index can be represented in the form of vectors. Since the vector dimensions of the user preference features and the feature index may be different, it is necessary to pad the dimension of the shorter vector in the user preference features and the feature index to be the same as that of the longer vector.
[0096] Calculating the cosine similarity between the user preference features and each feature index; in this step, the matching degree is measured by the cosine similarity, that is, calculating the cosine of the included angle between two vectors to measure the degree of similarity in the direction of the two vectors.
[0097] Obtaining the retrieval enhancement data corresponding to several feature indexes with the highest cosine similarity and integrating them into the second retrieval enhancement dataset.
[0098] In this embodiment, the user preference features and the feature index are processed by the padding method, so that this solution is not limited to the scenario of a fixed feature dimension; by calculating the cosine similarity between the user preference features and each feature index, several retrieval enhancement data that best match the user preference features can be found, thereby improving the relevance between the user preference features and the matched second retrieval enhancement dataset and providing a data basis for personalized content generation.
[0099] In one embodiment, please refer to Figure 5, the content creation module includes a feature extraction module, a feature screening module, and a content generation module; the feature extraction module is used to obtain the data features of the first retrieval-enhanced dataset and the second retrieval-enhanced dataset; the feature screening module is used to remove the parts of the data features of the first retrieval-enhanced dataset and the second retrieval-enhanced dataset that are not relevant to the standard intent data, and obtain an intent-related feature set; the content generation module encodes the standard intent data and the intent-related feature set into the same vector space through a Transformer encoder, and generates content data of the corresponding modality and its corresponding content descriptor vector through a number of integrated generation models.
[0100] Based on the foregoing embodiments, the first retrieval-enhanced dataset and the second retrieval-enhanced dataset can be retrieval-enhanced data of different modalities; the feature extraction module can perform feature extraction on the first retrieval-enhanced dataset and the second retrieval-enhanced dataset through a preset model. For example, image feature data is extracted through a convolutional neural network, image feature data of a video is extracted through a 3D convolutional neural network, text feature data is extracted through models such as Word2Vec and GPT, and audio feature data is extracted by combining Mel-frequency cepstral coefficients with a convolutional neural network. The feature screening module, based on the features extracted by the foregoing feature extraction module, retains the features related to the standard intent data through alignment analysis and correlation measurement, and obtains an intent-related feature set. The content generation module independently encodes the features in the standard intent data and the intent-related feature set based on a Transformer encoder, and then fuses the encoding results into the same vector space; in the same vector space, the target modality of the content data is selected according to the standard intent data, the preset generation model is determined according to the target modality, and then the content data of the corresponding modality and its corresponding content descriptor vector are generated through the generation model. Among them, the target modality can be one or more, and each target modality corresponds to a generation model.
[0101] In this embodiment, the data features of the retrieval-enhanced data are obtained through the feature extraction module, and the data features that are not relevant to the standard intent data are removed through the feature screening module, so as to reduce noise and irrelevant features, and at the same time improve the model's ability to accurately capture and understand the intent; through the content generation module, the features are encoded into the same vector space, and then the content data of the corresponding modality and its corresponding content descriptor vector are generated based on a number of integrated generation models, realizing efficient, accurate, and unified multi-modal content data generation, and ensuring the quality and accuracy of the content data.
[0102] In the embodiments provided in the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the modules can be in electrical, mechanical or other forms.
[0103] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0104] In addition, in each embodiment of the present application, the functional modules can be integrated into a processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules.
[0105] If the above-mentioned integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, read-only memory), random access memories (RAM, random access memory), magnetic disks, or optical discs that can store program codes.
[0106] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.
Claims
1. A content creation and multi-modal display platform based on AIGC, characterized in that: It includes an interaction management module, a user behavior analysis module, and a content creation module that are communicatively connected; The interaction management module is used to obtain a standard instruction word vector and a first retrieval enhancement dataset according to text or voice data, and generate standard intent data according to the standard instruction word vector and context information; The user behavior analysis module is used to obtain a second retrieval enhancement dataset according to the historical content description word vector of the user; The content creation module is used to generate multimodal content data and its corresponding content description word vector according to the standard intent data, the first retrieval enhancement dataset, and the second retrieval enhancement dataset; The interaction management module includes an input processing module, a context management module, and an intent recognition module; the input processing module is used to process text or voice data into a standard instruction word vector and a first retrieval enhancement dataset; The context management module is used to generate context information according to the dialogue history instruction word vector and the content description word vector; The intent recognition module is used to generate standard intent data according to the standard instruction word vector and context information; Generating the standard intent data according to the standard instruction word vector and context information includes the steps of: Integrating the standard instruction word vector and context information into a number of continuous word vector sequences and word vector type sequences according to the data timestamp; the word vector sequence consists of a start identifier, an end identifier, and a number of token IDs; Inputting the word vector sequence and the word vector type sequence into a number of layers of Transformer to generate the context representation of each token ID; extracting the start identifier from the last layer of Transformer as the comprehensive feature of the word vector sequence; Outputting the intent recognition result with the highest probability through a fully connected layer and a Softmax layer as the standard intent data.
2. The content creation and multi-modal display platform based on AIGC according to claim 1, characterized in that: The input processing module includes an instruction input module and a retrieval enhancement input module; the instruction input module is used to obtain the original operation instruction according to text or voice data and process the original operation instruction into a standard instruction word vector; The retrieval enhancement input module is used to obtain a first retrieval enhancement dataset according to text or voice data, construct the retrieval keywords corresponding to the first retrieval enhancement dataset, and store them.
3. The AIGC-based content creation and multi-modal display platform according to claim 2, characterized in that: The original operation instruction is a text instruction or a voice instruction; the instruction input module includes a speech-to-text conversion module and a first word vector extraction module; the speech-to-text conversion module is used to convert the voice instruction into a text instruction through a modality conversion strategy when the original operation instruction is a voice instruction; the first word vector extraction module is used to process the text instruction through a word vector extraction strategy to obtain the standard instruction word vector of the original operation instruction.
4. The AIGC-based content creation and multi-modal display platform according to claim 3, characterized in that: The modality conversion strategy specifically includes: Dividing the voice instruction into a number of audio frames, and preprocessing each audio frame to obtain the corresponding spectral data; Obtaining Mel spectral data through a Mel filter, performing cepstrum analysis based on the Mel spectral data to obtain Mel frequency cepstral coefficients, and integrating the Mel frequency cepstral coefficients into the current audio feature; Converting the current audio feature into a text instruction through a speech recognition model.
5. The AIGC-based content creation and multi-modal display platform according to claim 4, characterized in that: The Mel spectral data is expressed as: , , Among them, represents the m-th Mel spectrum data, where m is the number of Mel filters; k represents the number of frequency points, and N represents the total number of frequency points; represents the power spectrum of the k-th frequency point of the frequency-domain signal; represents the value of the m-th Mel filter at the frequency point k; represents the central position of the m-th Mel filter on the linear spectrum; The Mel frequency cepstral coefficients are expressed as: , Among them, M represents the number of Mel filters; n represents the serial number of the Mel-frequency cepstral coefficients; represents the value of the nth Mel-frequency cepstral coefficient.
6. The content creation and multi-modal display platform based on AIGC according to claim 3, characterized in that: The word vector extraction strategy specifically includes: Using a word segmentation algorithm to process the text instruction to obtain a number of instruction word segments; Representing and preprocessing the number of instruction word segments and inputting them into the Skip-gram model to obtain the standard instruction word vector of the original operation instruction.
7. The AIGC-based content creation and multi-modal display platform according to claim 1, characterized in that: The user behavior analysis module includes an intention data management module, a portrait analysis module, and a data matching module; the intention data management module is used to store the standard intention data and generate the historical intention data of the user according to the standard intention data; the portrait analysis module is used to generate user preference features according to the historical content description word vector of the user; the data matching module is used to store the retrieved enhanced data and its corresponding feature index, and obtain the second retrieved enhanced data set by calculating the matching degree between the user preference features and the feature index.
8. The AIGC-based content creation and multi-modal display platform according to claim 7, characterized in that: The portrait analysis module stores a preference word vector set, and the preference word vector set stores a number of preference word vectors; generating user preference features according to the historical content description word vector of the user specifically means generating user preference features according to the preference word vector set and the historical content description word vector of the user; the user preference features store a number of groups of preference word vectors and their corresponding frequencies.
9. The content creation and multi-modal display platform based on AIGC according to claim 1, characterized in that: The content creation module includes a feature extraction module, a feature screening module, and a content generation module; the feature extraction module is used to obtain the data features of the first retrieved enhanced data set and the second retrieved enhanced data set; The feature screening module is used to remove the parts of the data features of the first retrieved enhanced data set and the second retrieved enhanced data set that are not relevant to the standard intention data to obtain an intention-related feature set; The content generation module encodes the standard intention data and the intention-related feature set into the same vector space through a Transformer encoder, and generates content data of the corresponding modality and its corresponding content description word vector through a number of integrated generation models.
Citation Information
Patent Citations
Database-based retrieval enhancement and question and answer method and system
CN118364087A
Composite symbolic and non-symbolic artificial intelligence system for advanced reasoning and semantic search
US20240386015A1