Content data generation method and device based on artificial intelligence, equipment and medium
By employing a multimodal fusion technique based on hierarchical composite attention, the problem of low efficiency in existing multimodal fusion methods is solved, enabling efficient video content description generation and improving the marketing effectiveness of financial products.
Patent Information
- Application Number
- CN202510681937.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-10-28
AI Technical Summary
Existing multimodal fusion methods cannot efficiently integrate and utilize information from different modalities when generating short video content descriptions, resulting in low generation efficiency and difficulty in meeting the needs for fast and accurate descriptions in practical applications. In particular, they cannot accurately capture key information in short video marketing of financial products, thus reducing marketing effectiveness.
Employing a multimodal fusion technique based on hierarchical composite attention, this technique extracts, fuses, and generates visual, audio, and textual features through an input layer, a hierarchical composite attention module, and a decision module. It includes an intramodal attention layer, a cross-modal sparse attention layer, and a global-local refinement layer, reducing computational complexity while maintaining high-quality multimodal representation capabilities.
It effectively improves the generation efficiency of video content descriptions, reducing computational complexity from quadratic to approximately linear, thereby enhancing information dissemination and marketing effectiveness.
Smart Images

Figure CN120849658A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology and can be applied to fields such as fintech and digital healthcare, particularly to methods, devices, computer equipment, and storage media for generating content data based on artificial intelligence. Background Technology
[0002] In the fields of human-computer interaction and computer vision, multimodal fusion technology has always been a hot research topic. However, current multimodal fusion methods have significant drawbacks. The computational complexity of cross-attention increases quadratically with the number of modalities, which greatly limits the efficiency of models in large-scale multimodal data processing.
[0003] In practical applications such as short video content description, existing methods are particularly weak in handling multimodal temporal alignment and cross-modal sparse interactions. Multimodal temporal alignment ensures accurate correspondence of information from different modalities over time, while cross-modal sparse interactions help extract key and effective correlation information between different modalities. The shortcomings of existing methods result in an inefficient integration and utilization of information from different modalities when generating short video content descriptions, leading to low generation efficiency and failing to meet the needs of fast and accurate descriptions in practical applications.
[0004] For example, in the context of short video marketing content descriptions for financial products in the financial sector, these videos typically contain multimodal information, including visuals (such as product demonstrations and charts), text (such as speech-to-text product introductions and subtitles), and audio (such as background music and narration). Existing methods, due to poor time alignment, may associate key visuals demonstrating product benefits with incorrect narration segments, leading to discrepancies between the description and the actual product information. Furthermore, insufficient cross-modal sparse interaction processing fails to accurately capture effective connections between different modalities regarding key information such as product risks and returns, making it difficult for the generated marketing content to highlight product advantages, attract potential customers, and reduce the marketing effectiveness of financial products. Therefore, there is an urgent need for an efficient multimodal fusion method to improve the generation efficiency of short video content descriptions, thereby enhancing information dissemination and marketing effectiveness in various application scenarios, including financial products. Summary of the Invention
[0005] The purpose of this application is to propose a content data generation method, apparatus, computer device, and storage medium based on artificial intelligence, in order to solve the technical problem that when using existing multimodal fusion technology to generate short video content descriptions, it is impossible to efficiently integrate and utilize information from different modalities, resulting in low efficiency in content description generation.
[0006] Firstly, an artificial intelligence-based content data generation method is provided, including:
[0007] Receives visual input, audio input, and text input corresponding to the target video;
[0008] A preset fusion processing model is invoked; wherein, the fusion processing model includes an input layer, a hierarchical composite attention module, and a decision module;
[0009] Based on the input layer, feature extraction processing is performed on the visual input, the audio input, and the text input to obtain corresponding visual features, audio features, and text features;
[0010] The hierarchical composite attention module performs feature fusion processing on the visual features, audio features, and text features to obtain corresponding fused features; wherein, the hierarchical composite attention module includes an intramodal attention layer, a cross-modal sparse attention layer, and a global-local refinement layer;
[0011] The decision module processes the fusion features to generate a content description corresponding to the target video.
[0012] The content description is then processed for output.
[0013] Secondly, an artificial intelligence-based content data generation device is provided, comprising:
[0014] The receiving module is used to receive visual input, audio input, and text input corresponding to the target video.
[0015] The first invocation module is used to invoke a preset fusion processing model; wherein, the fusion processing model includes an input layer, a hierarchical composite attention module, and a decision module;
[0016] The extraction module is used to perform feature extraction processing on the visual input, the audio input and the text input based on the input layer to obtain the corresponding visual features, audio features and text features;
[0017] The fusion module is used to perform feature fusion processing on the visual features, the audio features, and the text features based on the hierarchical composite attention module to obtain the corresponding fused features; wherein, the hierarchical composite attention module includes an intramodal attention layer, a cross-modal sparse attention layer, and a global-local refinement layer;
[0018] The processing module is used to process the fusion features based on the decision module to generate a content description corresponding to the target video;
[0019] The output module is used to process the output of the content description.
[0020] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described artificial intelligence-based content data generation method.
[0021] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described content data generation method based on artificial intelligence.
[0022] In the aforementioned scheme implemented by the AI-based content data generation method, apparatus, computer equipment, and storage medium, visual input, audio input, and text input corresponding to the target video are first received; then, a preset fusion processing model is invoked; wherein, the fusion processing model includes an input layer, a hierarchical composite attention module, and a decision module; subsequently, feature extraction processing is performed on the visual input, audio input, and text input based on the input layer to obtain corresponding visual features, audio features, and text features; subsequently, feature fusion processing is performed on the visual features, audio features, and text features based on the hierarchical composite attention module to obtain corresponding fused features; wherein, the hierarchical composite attention module includes an intramodal attention layer, a cross-modal sparse attention layer, and a global-local refinement layer; further, the fused features are processed based on the decision module to generate a content description corresponding to the target video; finally, the content description is output. This application receives visual, audio, and text inputs corresponding to a target video. Then, based on a fusion processing model, it extracts features from the visual, audio, and text inputs through an input layer to obtain visual, audio, and text features. Next, a hierarchical composite attention module fuses these features to obtain fused features. A decision module then processes these fused features to generate a content description corresponding to the target video. Finally, the content description is output. Thus, this application employs a hierarchical composite attention-based multimodal fusion technology, using a three-layer attention architecture to achieve dynamic, efficient, and complete multimodal data fusion. This reduces computational complexity from quadratic to approximately linear while maintaining high-quality multimodal representation capabilities, effectively improving the generation efficiency of video content descriptions. Attached Figure Description
[0023] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is an exemplary system architecture diagram to which this application can be applied;
[0025] Figure 2 This is a flowchart of an embodiment of the content data generation method based on artificial intelligence according to this application;
[0026] Figure 3 This is a schematic diagram of a structure of an embodiment of the AI-based content data generation apparatus according to this application;
[0027] Figure 4 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation
[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.
[0029] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0030] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0031] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0032] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.
[0033] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.
[0034] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.
[0035] It should be noted that the content data generation method based on artificial intelligence provided in this application is generally executed by a server / terminal device, and correspondingly, the content data generation device based on artificial intelligence is generally set in the server / terminal device.
[0036] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0037] Continue to refer Figure 2 This document illustrates a flowchart of an embodiment of the AI-based content data generation method according to this application. The order of steps in the flowchart can be changed, and some steps can be omitted, depending on different needs. The AI-based content data generation method provided in this application can be applied to any scenario requiring content data generation, and thus can be applied to products in these scenarios, such as content data generation scenarios in the financial and medical fields. The AI-based content data generation method includes the following steps:
[0038] Step S201: Receive visual input, audio input, and text input corresponding to the target video.
[0039] In this embodiment, the content data generation method based on artificial intelligence runs on an electronic device (e.g., Figure 1 The server / terminal device shown can obtain visual input, audio input, and text input corresponding to the target video through wired or wireless connection. It should be noted that the above wireless connection methods may include, but are not limited to, 3G / 4G / 5G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other known or future wireless connection methods. The execution subject of this application is specifically a data generation system, which can be referred to as the system. This application can be applied to content data generation scenarios in the financial and medical fields. This application proposes a multimodal fusion technology based on hierarchical composite attention, which realizes dynamic, efficient and complete multimodal data fusion through a three-level attention architecture. The core innovations include: (1) modality-specific attention mechanism, which tailors the most suitable attention type for different modalities; (2) dynamic sparse cross attention, which selectively establishes intermodal connections through learnable gating; (3) global-local joint refinement, which balances macroscopic semantics and microscopic detail representation. This technology effectively addresses the limitations of existing fusion methods by reducing computational complexity from quadratic to approximately linear through hierarchical attention combination and sparse computation optimization, while maintaining high-quality multimodal representation capabilities.
[0040] The aforementioned target video is a short video input by the user that requires the generation of corresponding content descriptions. For example, in the financial field, the target video could be a video related to insurance alerts. The video content might include: The video begins with a young office worker busy at work when a large medical expense alert pops up on their phone, causing the worker to look worried. The scene then switches to the worker's home, where they are discussing the financial burden of this expense with their family, their expressions anxious. The video then shows them flipping through various insurance documents, looking confused and unsure how to choose the right insurance to address this risk. Finally, the scene freezes on the worker sighing helplessly. Alternatively, in the medical field, the target video could be a short video related to medical consultations. The video content might include: The video begins with an elderly person sitting on a living room sofa, clutching their chest, looking pained and slightly breathless. The scene then switches to the elderly person's children rushing home, their faces full of worry, comforting the elderly person while quickly packing to take them to the hospital. Upon arrival at the hospital, the scene shows registration and waiting, with the children anxiously looking around. Next, the doctor examines the elderly person, his expression focused, carefully inquiring about their symptoms. Finally, the scene shows the doctor giving a diagnosis and explaining the treatment plan in detail to the children, who listen attentively, nodding occasionally.
[0041] Step S202: Invoke the preset fusion processing model; wherein the fusion processing model includes an input layer, a hierarchical composite attention module, and a decision module.
[0042] In this embodiment, the construction process of the above-mentioned fusion processing model will be described in more detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0043] Step S203: Based on the input layer, perform feature extraction processing on the visual input, the audio input, and the text input to obtain the corresponding visual features, audio features, and text features.
[0044] In this embodiment, the specific implementation process of performing feature extraction processing on the visual input, audio input, and text input based on the input layer to obtain the corresponding visual features, audio features, and text features will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0045] Step S204: The aforementioned hierarchical composite attention module comprises three stages of processing: an intra-modal attention layer, a cross-modal sparse attention layer, and a global-local refinement layer. The hierarchical composite attention module outputs the final fused feature, which is then fed into the decision module to generate the final output. Specifically, the visual features, audio features, and text features are fused based on the hierarchical composite attention module to obtain the corresponding fused feature; wherein the hierarchical composite attention module includes an intra-modal attention layer, a cross-modal sparse attention layer, and a global-local refinement layer.
[0046] In this embodiment, the specific implementation process of performing feature fusion processing on the visual features, audio features, and text features based on the hierarchical composite attention module to obtain the corresponding fused features will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0047] Step S205: Process the fusion features based on the decision module to generate a content description corresponding to the target video.
[0048] In this embodiment, the main task of the decision module is to convert the fused features into a natural language description. The decision module consists of a decoder, which can be a Recurrent Neural Network (RNN) based decoder (such as LSTM or GRU) or a Transformer based decoder. During the inference phase, the process of generating the content description includes decoder initialization and decoding. Specifically, decoder initialization includes using the fused features as the decoder's initial state. Specifically, the initial hidden state can be mapped to the decoder's hidden state space through a fully connected layer. The decoding process includes word embedding: for each time step of generating the description, a starting token (such as...) is first needed. <start>The decoder input is initialized using a _________. Input processing: The input of the current time step (usually the word embedding vector generated in the previous time step) is fed into the decoder. Attention mechanism: At each time step of the decoder, an attention mechanism can be used to focus on different parts of the fused features to better capture contextual information. Predicting the next word: The decoder output is passed through a fully connected layer and a softmax layer to predict the probability distribution of the next word. Sampling or greedy search: Based on the predicted probability distribution, a greedy search (selecting the word with the highest probability) or a bundle search (keeping multiple candidate sequences) can be selected to generate a description. Termination condition. End marker: When the decoder generates an end marker (such as...), the end marker is used to terminate the decoding process. <end>When the generated content description reaches the specified length, the generation process stops. Maximum length limit: To avoid infinite loops, a maximum generation length can be set; generation will stop when the generated content description reaches this length.
[0049] Step S206: Output the content description.
[0050] In this embodiment, the specific implementation process of outputting the content description described above will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0051] This application first receives visual input, audio input, and text input corresponding to the target video; then it calls a preset fusion processing model; wherein the fusion processing model includes an input layer, a hierarchical composite attention module, and a decision module; then, based on the input layer, it performs feature extraction processing on the visual input, the audio input, and the text input to obtain corresponding visual features, audio features, and text features; subsequently, based on the hierarchical composite attention module, it performs feature fusion processing on the visual features, the audio features, and the text features to obtain corresponding fused features; wherein the hierarchical composite attention module includes an intramodal attention layer, a cross-modal sparse attention layer, and a global-local refinement layer; further, based on the decision module, it processes the fused features to generate a content description corresponding to the target video; finally, it outputs the content description. This application receives visual, audio, and text inputs corresponding to a target video. Then, based on a fusion processing model, it extracts features from the visual, audio, and text inputs through an input layer to obtain visual, audio, and text features. Next, a hierarchical composite attention module fuses these features to obtain fused features. A decision module then processes these fused features to generate a content description corresponding to the target video. Finally, the content description is output. Thus, this application employs a hierarchical composite attention-based multimodal fusion technology, using a three-layer attention architecture to achieve dynamic, efficient, and complete multimodal data fusion. This reduces computational complexity from quadratic to approximately linear while maintaining high-quality multimodal representation capabilities, effectively improving the generation efficiency of video content descriptions.
[0052] In some optional implementations of this embodiment, before step S202, the electronic device may further perform the following steps:
[0053] Obtain pre-generated sample data.
[0054] In this embodiment, a dataset containing matching visual, audio, and text data is pre-collected and used as the sample data. The sample data is then divided into multiple batches, each containing a certain number of samples. Simultaneously, a validation set and a test set are prepared to monitor model performance and evaluate the final results.
[0055] Invokes a pre-built initial model containing the specified structure.
[0056] In this embodiment, the initial model described above is a pre-constructed model with an input layer, a hierarchical composite attention module, and a decision module. First, all learnable parameters in the model can be initialized, including the projection matrix, attention weights (implicitly generated during attention calculation), gating parameters, and optimizer-related parameters (such as the momentum of the Adam optimizer). The initialization method can employ common random initialization techniques, such as random sampling from a Gaussian distribution.
[0057] Obtain the preset multi-task loss function.
[0058] In this embodiment, the multi-task loss function includes:
[0059]
[0060] in represents the task-specific loss function; ||M||1 represents the L1 norm of the interaction matrix (||·||1 represents the L1 norm of a vector or matrix, i.e., the sum of the absolute values of all elements), used to constrain sparsity; λ1, λ2, and λ3 represent the orthogonality loss, which promotes complementarity between modes; λ1, λ2, and λ3 are the weighting coefficients that balance different loss terms.
[0061] Based on the preset end-to-end method and the multi-task loss function, the initial model is trained using the sample data to obtain the trained specified model.
[0062] In this embodiment, the initial model is trained in an end-to-end manner and through multi-task loss joint optimization, using sample data to train the initial model, thereby obtaining the trained specified model, which is then used as the required fusion processing model.
[0063] The specified model is used as the fusion processing model.
[0064] This application obtains pre-generated sample data; then calls a pre-built initial model containing a specified structure; subsequently obtains a preset multi-task loss function; and then trains the initial model using the sample data based on a preset end-to-end approach and the multi-task loss function to obtain a trained specified model; finally, the specified model is used as the fusion processing model. This application, by obtaining pre-generated sample data, calling a pre-built initial model containing a specified structure, obtaining a preset multi-task loss function, and then training the initial model using the sample data based on a combination of end-to-end approach and the multi-task loss function, can efficiently and accurately construct the required fusion processing model, improving the model construction efficiency of the fusion processing model and ensuring the model processing effect of the obtained fusion processing model.
[0065] In some alternative implementations, step S204 includes the following steps:
[0066] Based on the intramodal attention layer in the hierarchical composite attention module, modal processing is performed on the visual features, audio features, and text features to obtain the corresponding target visual features, target audio features, and target text features.
[0067] In this embodiment, the specific implementation process of performing modal processing on the visual features, audio features, and text features based on the intramodal attention layer in the hierarchical composite attention module to obtain the corresponding target visual features, target audio features, and target text features will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0068] Based on the cross-modal sparse attention layer, sparse interaction processing is performed between different features of the target visual features, the target audio features, and the target text features to obtain a first cross-modal attention feature between the target visual features and the target audio features, a second cross-modal attention feature between the target visual features and the target text features, and a third cross-modal attention feature between the target audio features and the target text features.
[0069] In this embodiment, the cross-modal sparse attention layer aims to explore the correlations between different modalities. By generating a sparse interaction matrix and calculating cross-modal attention based on this matrix, the interaction and fusion between features of different modalities are realized. For example, the attention Cv,a of the visual-audio modality, the attention Cv,t of the visual-text modality, and the attention Ca,t of the audio-text modality are calculated. These cross-modal attention results contain important correlation information between different modalities.
[0070] Specifically, to reduce computational complexity and preserve key cross-modal interactions, a sparse interaction matrix M∈[0,1] is predefined. 3×3 (M represents the modal interaction matrix, M) i,j =1 indicates that mode i is allowed to interact with mode j. This matrix is generated in two steps:
[0071] Static predefined: Setting up basic interaction pairs based on prior knowledge, such as M v,a =1 (Visual-audio is effective for action scenes), M a,t =1 (audio-text is effective for speech transcription), M v,t =0 (Direct interaction can be disabled when an audio intermediary is present).
[0072] Dynamic adjustment: Corrected weights ΔM∈[0,1] are generated through a gating network. 3×3 (ΔM represents the dynamically adjusted weight matrix):
[0073] ΔM=σ(MLP([V ′ A ′ ;T ′ ]))
[0074] MLP stands for Multilayer Perceptron, a type of feedforward neural network; [V] ′ A ′ ;T ′ The ] symbol represents the concatenated modal features. The final cross-modal attention calculation is as follows:
[0075]
[0076] The Sparsemax function projects the attention weights onto a probability simplex, resulting in a sparse distribution; C m,n This represents the interaction characteristics between modes $m$ and $n$. and It is the projection matrix; m and n represent the features of different modes, respectively.
[0077] Based on the global-local refinement layer, feature fusion processing is performed on the first cross-modal attention feature, the second cross-modal attention feature, and the third cross-modal attention feature to obtain the corresponding processed features.
[0078] Furthermore, the cross-modal attention results Cv,a,Cv,t, and Ca,t obtained from the cross-modal attention layer are concatenated and used as input to the global-local refinement layer. These features, which incorporate multimodal information, provide rich contextual information for subsequent global and local attention calculations, enabling the model to refine features from a more comprehensive perspective.
[0079] In this embodiment, the specific implementation process of performing feature fusion processing on the first cross-modal attention feature, the second cross-modal attention feature, and the third cross-modal attention feature based on the global-local refinement layer to obtain the corresponding processed features will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0080] The processing features are used as the fusion features.
[0081] This application performs modal processing on the visual features, audio features, and text features based on the intramodal attention layer in the hierarchical composite attention module to obtain corresponding target visual features, target audio features, and target text features. Then, based on the cross-modal sparse attention layer, it performs sparse interaction processing between different features on the target visual features, target audio features, and target text features to obtain a first cross-modal attention feature between the target visual features and the target audio features, a second cross-modal attention feature between the target visual features and the target text features, and a third cross-modal attention feature between the target audio features and the target text features. Subsequently, based on the global-local refinement layer, it performs feature fusion processing on the first cross-modal attention feature, the second cross-modal attention feature, and the third cross-modal attention feature to obtain corresponding processed features. Finally, the processed features are used as the fused features. This application, through the collaboration of the intramodal attention layer, cross-modal sparse attention layer, and global-local refinement layer contained in the hierarchical composite attention module, can achieve the process of gradually completing the process from intramodal feature extraction to cross-modal interaction, and then to feature refinement and optimization, thereby improving the fusion processing model's ability to process multimodal data and ensuring the accuracy of the generated fusion features.
[0082] In some optional implementations of this embodiment, the modal processing of the visual features, audio features, and text features based on the intramodal attention layer in the hierarchical composite attention module to obtain the corresponding target visual features, target audio features, and target text features includes the following steps:
[0083] Within the intramodal attention layer, the visual features are processed using a preset scaled dot product attention method to obtain the target visual features corresponding to the visual features.
[0084] In this embodiment, the intramodal attention layer employs a modality-specific attention mechanism to select the most suitable attention type for data with different characteristics. Specifically, visual modality processing includes: using scaled dot product attention and introducing a learnable temperature parameter τ. v (τ v This represents a visual attention temperature parameter, controlling the smoothness of attention distribution.
[0085]
[0086] Wherein, Softmax represents the normalized activation function, which transforms the input vector into a probability distribution; Represents the query matrix. Represents the key matrix. It is a learnable projection matrix, and \top denotes the matrix transpose operation. Temperature parameter τ v Initialized to 1.0, it is automatically learned through backpropagation to adjust the level of attentional focus on visual features. A smaller τ... v Values (such as 0.5-1.2) produce a sharper attention distribution, suitable for handling detail-sensitive tasks (such as hand gesture recognition). Furthermore, visual features can be processed based on the aforementioned visual modality processing steps to obtain target visual features corresponding to the visual features.
[0087] The audio features are processed using a preset additive attention mechanism to obtain target audio features corresponding to the audio features.
[0088] In this embodiment, the above-mentioned audio modality processing includes: designing a hybrid attention gating system that combines the advantages of dot product and additive attention.
[0089] A′=σ(g a )⊙A dot +(1-σ(g a ))⊙A add
[0090] Where σ represents the Sigmoid activation function, which maps the input to the (0,1) interval; ⊙ represents element-wise multiplication (Hadamard product); g a ∈R 512 It is a learnable gated vector; A dot A represents the result of dot product attention. add This represents the result of additive attention. The gating mechanism can adaptively select the optimal attention type, favoring dot-product attention for ambient noise and additive attention for speech content. Furthermore, the audio features can be processed based on the aforementioned audio modality processing steps to obtain the target audio features corresponding to those features.
[0091] Based on a preset hybrid attention gating, text modality processing is performed on the text features to obtain target text features corresponding to the text features.
[0092] In this embodiment, the above text modality processing includes: using an additive attention mechanism to define a learnable query vector q. t ∈R 512 (q t The query vector representing the text modality, used to guide attention calculation:
[0093]
[0094] Where tanh represents the hyperbolic tangent activation function, which compresses the input values to the interval [-1, 1]; v t ∈R 512 and W t ∈R 512×512 These are learnable parameters that collectively determine the attention weights of text features; T represents the attention weight of the i-th text feature; ′ This represents the weighted text representation. This mechanism is particularly well-suited for processing long texts (T). t (>100), which can automatically focus on key instruction words. In addition, the above text modality processing steps can be used to process the above text features to obtain the target audio features corresponding to the text features.
[0095] This application performs visual modality processing on the visual features based on a preset scaled dot product attention mechanism within the intramodal attention layer to obtain target visual features corresponding to the visual features; performs audio modality processing on the audio features based on a preset additive attention mechanism to obtain target audio features corresponding to the audio features; and performs text modality processing on the text features based on a preset hybrid attention gating to obtain target text features corresponding to the text features. By employing a modality-specific attention mechanism in the intramodal attention layer, this application can select the most suitable attention type for modal feature data with different characteristics for modality processing, thereby effectively improving the intelligence and adaptability of modality feature data processing and ensuring the accuracy of the obtained target visual features, target audio features, and target text features. Furthermore, by enhancing the intramodal feature representation, the intramodal attention layer provides high-quality data input for cross-modal sparse attention layers.
[0096] In some optional implementations, the global-local refinement layer includes a global attention branch and a local attention branch; the feature fusion processing of the first cross-modal attention feature, the second cross-modal attention feature, and the third cross-modal attention feature based on the global-local refinement layer to obtain the corresponding processed features includes the following steps:
[0097] Based on the global attention branch in the global-local refinement layer, global attention is calculated on the first cross-modal attention feature, the second cross-modal attention feature, and the third cross-modal attention feature to obtain the corresponding global features.
[0098] In this embodiment, to simultaneously capture both macroscopic semantics and microscopic details, the aforementioned global-local refinement layer includes two branches: a global attention branch and a local attention branch. The global attention calculation includes: using the global attention branch to calculate attention across the entire sequence.
[0099]
[0100] Where F = [C v,a C v,t C a,t [] indicates the concatenated cross-modal features; It is the projection matrix; F g This represents the features after global attention processing (i.e., global features). This branch captures the overall semantic structure of the video, such as the complete recognition process.
[0101] Based on the local attention branch, local attention calculations are performed on the first cross-modal attention feature, the second cross-modal attention feature, and the third cross-modal attention feature to obtain the corresponding local features.
[0102] In this embodiment, the above-mentioned local attention calculation includes: processing using a sliding window (window size w = 8) through local attention branches:
[0103]
[0104] Among them, F i:i+w This represents the feature subsequence from position i to i+w; W represents the attention result (i.e., local features) within a local window; l Q and W l K It is a projection matrix. Each window independently calculates attention, focusing on local temporal relationships, such as accurately capturing continuous actions within a short period of time.
[0105] The global features and the local features are connected based on residual connections to obtain the corresponding specified features.
[0106] In this embodiment, global and local features can be merged through residual connections:
[0107] F ′ =LayerNorm(F g +F l )
[0108] Where LayerNorm represents the layer normalization operation, used to stabilize the training process; F ′ This represents the final fused features. This design preserves both macroscopic context and microscopic details. Experiments show that the joint strategy significantly improves the performance of action segmentation tasks compared to using global attention alone.
[0109] The specified feature is used as the processing feature.
[0110] This application calculates global attention for the first, second, and third cross-modal attention features based on the global attention branch in the global-local refinement layer to obtain corresponding global features; and calculates local attention for the same features based on the local attention branch to obtain corresponding local features; then, it connects the global and local features using residual connections to obtain corresponding specified features; these specified features are then used as the processed features. This application further processes cross-modal fusion features from both global and local perspectives by using the global-local refinement layer. The global attention branch captures long-distance dependencies in the entire feature sequence, while the local attention branch focuses on the relationships between local features through a sliding window. This combination of global and local approaches allows for a more comprehensive mining of information within the features. Furthermore, by merging global and local features through residual connections, the features obtained from the previous layers are integrated and optimized, resulting in more robust and discriminative output features that better serve subsequent content generation tasks.
[0111] In some alternative implementations, step S206 includes the following steps:
[0112] Obtain the preset content verification strategy.
[0113] In this embodiment, the above content verification strategy is to verify the semantics of the content description to ensure that the semantics of the content description are correct and the logic is clear.
[0114] The content description is validated based on the content validation strategy.
[0115] In this embodiment, the verification process for the above content description can be performed based on the content verification strategy described above. That is, the semantics of the content description are verified to be correct and the logic is clear. Only when the semantics of the content description are detected to be correct and the logic is clear will the content description be determined to pass the verification. Otherwise, the content description is determined to fail the verification.
[0116] If the content description passes verification, the preset output method is obtained.
[0117] In this embodiment, the selection of the above output method is not specifically limited and can be determined according to the user's needs. For example, any one of the following methods can be used: email sending, interface display, or message sending. Here, the "user" refers to the user who inputs visual input, audio input, and text input corresponding to the target video.
[0118] The content description is output based on the aforementioned output method.
[0119] In this embodiment, the above content description can be sent to the above user according to the selected output method to complete the output processing of the above content description.
[0120] This application obtains a preset content verification strategy; then verifies the content description based on the content verification strategy; if the content description passes verification, a preset output method is obtained; subsequently, the content description is output based on the output method. After generating a content description corresponding to the target video, this application intelligently verifies the content description based on the content verification strategy, and automatically outputs the content description based on the obtained output method after detecting that the content description has passed verification. This effectively ensures the accuracy and standardization of the output content description, thereby improving the user experience.
[0121] In some optional implementations of this embodiment, step S203 includes the following steps:
[0122] Obtain the preset projection strategy.
[0123] In this embodiment, the projection strategy described above can specifically employ linear projection.
[0124] Determine the feature space of the target dimension.
[0125] In this embodiment, the feature space of the target dimension can specifically be a 512-dimensional feature space.
[0126] Based on the projection strategy, the visual input, audio input, and text input are unified to the feature space of the target dimension through the input layer, thereby obtaining visual features corresponding to the visual input, audio features corresponding to the audio input, and text features corresponding to the text input.
[0127] In this embodiment, linear projection can be used to unify the inputs of each modality into a 512-dimensional feature space to form standardized inputs, including standardized visual features, standardized audio features, and standardized text features.
[0128] For example, suppose the system receives input from three modalities: visual input ($V$ represents the visual feature matrix, R represents the real number field, 2048 represents the feature dimension, T v (representing visual sequence length), audio input (A represents the audio feature matrix, 128 represents the feature dimension, T) a (representing audio sequence length) and text input (T represents the text feature matrix, 768 represents the feature dimension, T) t (This represents the length of the text sequence). The system then uses linear projection to unify the features of each modality into a standard 512-dimensional feature space, forming a standardized input. ( Represents standardized visual features, Represents standardized audio features, (This represents the standardized text features, where T represents the sequence length of the corresponding modality).
[0129] This application obtains a preset projection strategy; then determines the feature space of the target dimension; subsequently, based on the projection strategy, the visual input, audio input, and text input are unified to the feature space of the target dimension through the input layer, obtaining visual features corresponding to the visual input, audio features corresponding to the audio input, and text features corresponding to the text input. By using the obtained projection strategy and unifying visual, audio, and text inputs to the feature space of the target dimension through the input layer, this application can automatically and accurately obtain visual features corresponding to visual input, audio features corresponding to audio input, and text features corresponding to text input, ensuring the accuracy, standardization, and normalization of the extracted visual, audio, and text features.
[0130] In some alternative implementations, the user information obtained is subject to user consent and complies with relevant laws and policies.
[0131] Furthermore, any software tools or components not belonging to our company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.
[0132] Furthermore, this application has the following advantages:
[0133] (1) By using modality-specific temperature parameters and a learnable gating mechanism, customized processing of different modal features is achieved, which can automatically adjust the attention distribution according to task requirements, significantly improving the ability to capture key actions and semantics. The gating network response time is less than 2 milliseconds, ensuring that the system can operate efficiently in real-time scenarios.
[0134] (2) The combination of sparse attention matrix and dynamic gating network reduces the computational complexity of cross-modal interaction from O(n^2) to approximately O(nlogn) while retaining more than 90% of the key interactions. In a 512-dimensional feature space, GPU memory usage is reduced by 63%, enabling this technology to achieve real-time processing on edge devices and meet the needs of resource-constrained scenarios.
[0135] (3) The parallel design of the global-local refinement layer not only preserves macro-contextual information but also enhances the expressive power of fine-grained features. From the complete process of "preparing ingredients-cooking-plating" to the precise action sequence of "pouring oil-frying and adding salt", the system can accurately capture the content description BLEU-4 index by 17.2% on the MSR-VTT dataset.
[0136] (4) The layered composite attention module supports two integration modes: the replacement mode can be directly embedded into the existing system, and the plug-in mode can be used as an independent processor. This modular design enables the technology to quickly adapt to application scenarios in different industries, achieving seamless integration from industrial quality inspection to medical image analysis, truly achieving plug-and-play functionality.
[0137] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0138] It should be emphasized that, to further ensure the privacy and security of the above content description, the above content description can also be stored in a blockchain node.
[0139] The blockchain referred to in this application is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0140] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0141] Foundational technologies in artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0142] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware with computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When executed, the program can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).
[0143] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0144] Further reference Figure 3 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of an artificial intelligence-based content data generation device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0145] like Figure 3 As shown, the content data generation device 300 based on artificial intelligence described in this embodiment includes: a receiving module 301, a first calling module 302, an extraction module 303, a fusion module 304, a processing module 305, and an output module 306. Wherein:
[0146] The receiving module 301 is used to receive visual input, audio input, and text input corresponding to the target video;
[0147] The first invocation module 302 is used to invoke a preset fusion processing model; wherein, the fusion processing model includes an input layer, a hierarchical composite attention module, and a decision module;
[0148] Extraction module 303 is used to perform feature extraction processing on the visual input, the audio input and the text input based on the input layer to obtain corresponding visual features, audio features and text features;
[0149] The fusion module 304 is used to perform feature fusion processing on the visual features, the audio features, and the text features based on the hierarchical composite attention module to obtain the corresponding fused features; wherein, the hierarchical composite attention module includes an intramodal attention layer, a cross-modal sparse attention layer, and a global-local refinement layer.
[0150] Processing module 305 is used to process the fusion features based on the decision module to generate a content description corresponding to the target video;
[0151] The output module 306 is used to output the content description.
[0152] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the content data generation method based on artificial intelligence in the aforementioned implementation method, and will not be repeated here.
[0153] In some optional implementations of this embodiment, the content data generation device based on artificial intelligence further includes:
[0154] The first acquisition module is used to acquire pre-generated sample data;
[0155] The second calling module is used to call a pre-built initial model containing a specified structure;
[0156] The second acquisition module is used to acquire a preset multi-task loss function;
[0157] The training module is used to train the initial model using the sample data based on a preset end-to-end method and the multi-task loss function, so as to obtain a trained specified model.
[0158] The determination module is used to select the specified model as the fusion processing model.
[0159] In some optional implementations of this embodiment, the processing module 305 includes:
[0160] The first processing submodule is used to perform modal processing on the visual features, the audio features and the text features based on the intramodal attention layer in the hierarchical composite attention module to obtain the corresponding target visual features, target audio features and target text features.
[0161] The second processing submodule is used to perform sparse interaction processing between different features on the target visual features, the target audio features, and the target text features based on the cross-modal sparse attention layer, to obtain a first cross-modal attention feature between the target visual features and the target audio features, a second cross-modal attention feature between the target visual features and the target text features, and a third cross-modal attention feature between the target audio features and the target text features;
[0162] The third processing submodule is used to perform feature fusion processing on the first cross-modal attention feature, the second cross-modal attention feature, and the third cross-modal attention feature based on the global-local refinement layer to obtain the corresponding processed features;
[0163] The first determining submodule is used to use the processed features as the fused features.
[0164] In some optional implementations of this embodiment, the first processing submodule includes:
[0165] The first processing unit is used to perform visual modality processing on the visual features based on a preset scaled dot product attention within the modal attention layer to obtain target visual features corresponding to the visual features.
[0166] The second processing unit is used to perform audio modality processing on the audio features based on a preset additive attention mechanism to obtain target audio features corresponding to the audio features;
[0167] The third processing unit is used to perform text modality processing on the text features based on a preset hybrid attention gating to obtain target text features corresponding to the text features.
[0168] In some optional implementations of this embodiment, the global-local refinement layer includes a global attention branch and a local attention branch; the third processing submodule includes:
[0169] The first computing unit is used to perform global attention calculation on the first cross-modal attention feature, the second cross-modal attention feature, and the third cross-modal attention feature based on the global attention branch in the global-local refinement layer, so as to obtain the corresponding global features;
[0170] The second calculation unit is used to perform local attention calculations on the first cross-modal attention feature, the second cross-modal attention feature, and the third cross-modal attention feature based on the local attention branch, so as to obtain the corresponding local features;
[0171] A connection unit is used to perform connection processing on the global feature and the local feature based on residual connection to obtain the corresponding specified feature;
[0172] A determining unit is used to use the specified feature as the processing feature.
[0173] In some optional implementations of this embodiment, the processing module 305 includes:
[0174] The first acquisition submodule is used to acquire preset content verification strategies;
[0175] The verification submodule is used to verify the content description based on the content verification strategy;
[0176] The second acquisition submodule is used to acquire a preset output method if the content description passes the verification.
[0177] The output submodule is used to perform output processing on the content description based on the output method.
[0178] In some optional implementations of this embodiment, the extraction module 303 includes:
[0179] The third acquisition submodule is used to acquire the preset projection strategy;
[0180] The second determination submodule is used to determine the feature space of the target dimension;
[0181] The projection submodule is used to unify the visual input, the audio input, and the text input to the feature space of the target dimension through the input layer based on the projection strategy, so as to obtain the visual features corresponding to the visual input, the audio features corresponding to the audio input, and the text features corresponding to the text input.
[0182] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.
[0183] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected via a system bus. It should be noted that only the computer device 4 with components 41-43 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0184] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.
[0185] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 4. Of course, the memory 41 may also include both the internal storage unit and its external storage device of the computer device 4. In this embodiment, the memory 41 is typically used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for content data generation methods based on artificial intelligence. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or will be output.
[0186] In some embodiments, the processor 42 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 42 is typically used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions stored in the memory 41 or to process data, for example, to execute computer-readable instructions of the AI-based content data generation method.
[0187] The network interface 43 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 4 and other electronic devices.
[0188] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the above-described artificial intelligence-based content data generation method.
[0189] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0190] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.< / end> < / start>
Claims
1. A content data generation method based on artificial intelligence, characterized in that, Includes the following steps: Receives visual input, audio input, and text input corresponding to the target video; A preset fusion processing model is invoked; wherein, the fusion processing model includes an input layer, a hierarchical composite attention module, and a decision module; Based on the input layer, feature extraction processing is performed on the visual input, the audio input, and the text input to obtain corresponding visual features, audio features, and text features; The hierarchical composite attention module performs feature fusion processing on the visual features, audio features, and text features to obtain corresponding fused features; wherein, the hierarchical composite attention module includes an intramodal attention layer, a cross-modal sparse attention layer, and a global-local refinement layer; The decision module processes the fusion features to generate a content description corresponding to the target video. The content description is then processed for output.
2. The content data generation method based on artificial intelligence according to claim 1, characterized in that, Before the step of invoking the preset fusion processing model, the following is also included: Obtain pre-generated sample data; Invokes a pre-built initial model containing the specified structure; Obtain the preset multi-task loss function; Based on the preset end-to-end method and the multi-task loss function, the initial model is trained using the sample data to obtain the trained specified model; The specified model is used as the fusion processing model.
3. The content data generation method based on artificial intelligence according to claim 1, characterized in that, The step of performing feature fusion processing on the visual features, audio features, and text features based on the hierarchical composite attention module to obtain the corresponding fused features specifically includes: Based on the intramodal attention layer in the hierarchical composite attention module, modal processing is performed on the visual features, the audio features, and the text features to obtain the corresponding target visual features, target audio features, and target text features. Based on the cross-modal sparse attention layer, sparse interaction processing is performed between different features of the target visual features, the target audio features, and the target text features to obtain a first cross-modal attention feature between the target visual features and the target audio features, a second cross-modal attention feature between the target visual features and the target text features, and a third cross-modal attention feature between the target audio features and the target text features; Based on the global-local refinement layer, feature fusion processing is performed on the first cross-modal attention feature, the second cross-modal attention feature, and the third cross-modal attention feature to obtain the corresponding processed features; The processing features are used as the fusion features.
4. The content data generation method based on artificial intelligence according to claim 3, characterized in that, The step of performing modal processing on the visual features, audio features, and text features based on the intramodal attention layer in the hierarchical composite attention module to obtain the corresponding target visual features, target audio features, and target text features specifically includes: Within the intramodal attention layer, the visual features are subjected to visual modal processing based on a preset scaled dot product attention to obtain the target visual features corresponding to the visual features. The audio features are processed using a preset additive attention mechanism to obtain target audio features corresponding to the audio features. Based on a preset hybrid attention gating, text modality processing is performed on the text features to obtain target text features corresponding to the text features.
5. The content data generation method based on artificial intelligence according to claim 3, characterized in that, The global-local refinement layer includes a global attention branch and a local attention branch; the step of performing feature fusion processing on the first cross-modal attention feature, the second cross-modal attention feature, and the third cross-modal attention feature based on the global-local refinement layer to obtain the corresponding processed features specifically includes: Based on the global attention branch in the global-local refinement layer, global attention is calculated on the first cross-modal attention feature, the second cross-modal attention feature, and the third cross-modal attention feature to obtain the corresponding global features; Based on the local attention branch, local attention calculation is performed on the first cross-modal attention feature, the second cross-modal attention feature, and the third cross-modal attention feature to obtain the corresponding local features; The global features and the local features are connected based on residual connections to obtain the corresponding specified features; The specified feature is used as the processing feature.
6. The content data generation method based on artificial intelligence according to claim 1, characterized in that, The step of outputting the content description specifically includes: Obtain the preset content verification strategy; The content description is validated based on the content validation strategy. If the content description passes verification, the preset output method is obtained; The content description is output based on the aforementioned output method.
7. The content data generation method based on artificial intelligence according to claim 1, characterized in that, The step of performing feature extraction processing on the visual input, audio input, and text input based on the input layer to obtain corresponding visual features, audio features, and text features specifically includes: Obtain the preset projection strategy; Determine the feature space of the target dimension; Based on the projection strategy, the visual input, audio input, and text input are unified to the feature space of the target dimension through the input layer, thereby obtaining visual features corresponding to the visual input, audio features corresponding to the audio input, and text features corresponding to the text input.
8. A content data generation device based on artificial intelligence, characterized in that, include: The receiving module is used to receive visual input, audio input, and text input corresponding to the target video. The first invocation module is used to invoke a preset fusion processing model; wherein, the fusion processing model includes an input layer, a hierarchical composite attention module, and a decision module; The extraction module is used to perform feature extraction processing on the visual input, the audio input and the text input based on the input layer to obtain the corresponding visual features, audio features and text features; The fusion module is used to perform feature fusion processing on the visual features, the audio features, and the text features based on the hierarchical composite attention module to obtain the corresponding fused features; wherein, the hierarchical composite attention module includes an intramodal attention layer, a cross-modal sparse attention layer, and a global-local refinement layer; The processing module is used to process the fusion features based on the decision module to generate a content description corresponding to the target video; The output module is used to process the output of the content description.
9. A computer device, characterized in that, The method includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the content data generation method based on artificial intelligence as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the content data generation method based on artificial intelligence as described in any one of claims 1 to 7.