Sound asset management method, system and storage medium based on big data
By building a sound asset graph structure through big data analysis and AI technology, and using graph neural networks and generative adversarial networks to generate customized audio content, we solve the problem of intelligent material integration and generation in sound asset management, and achieve efficient and personalized audio content generation.
Patent Information
- Application Number
- CN202411279178.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-12
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2044-09-12
AI Technical Summary
The diversity of audio materials in existing sound asset management is difficult to effectively integrate, the reconstruction and generation technology has a low level of intelligence, and it is impossible to generate highly customized audio content, and it cannot meet users' needs for personalized and emotional sound assets.
Through big data analysis, graph neural networks and generative adversarial networks, we build a sound asset graph structure, screen matching materials based on semantic similarity and emotional relevance, use graph neural networks for feature fusion, and generate scene-derived sound assets that match the target emotional style through generative adversarial networks.
It significantly improves the intelligence level of sound asset management, can quickly screen out materials that meet specific emotional and semantic requirements, generate highly customized audio content, and enhance the application value and personalized expression capabilities of sound assets.
Smart Images

Figure CN119202304B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of big data mining and analysis, and in particular to a sound asset management method, system and storage medium based on big data. Background Art
[0002] With the rapid development of digital technology, especially in the context of building an international trading platform for sound assets and the National Soundscape Port construction project, the management and application of sound assets have been given more attention. In particular, the National Soundscape Port construction project can effectively assist the digital protection, innovation and transaction flow of national sound assets, and build a national sound asset full property rights ecosystem.
[0003] Currently, the primary challenge facing sound asset management lies in the sheer diversity of audio assets. Different audio assets exhibit significant differences in content, style, and emotional expression, making these differences difficult to effectively integrate and leverage through traditional management methods. Furthermore, current sound asset reconstruction and generation technologies lack intelligence, typically supporting only simple audio splicing or pitch shifting. This results in inefficient utilization of sound assets across diverse application scenarios, making it difficult to generate highly customized audio content and failing to meet user demands for personalized and emotionally engaging sound assets.
[0004] To address the above issues, the industry has not yet proposed a better technical solution. Summary of the Invention
[0005] The present application provides a sound asset management method, device, storage medium, computer program product and electronic device based on big data, which is used to at least solve the problem that the current sound asset reconstruction and generation technology has a low level of intelligence and is unable to support highly customized audio application scenarios.
[0006] In a first aspect, an embodiment of the present application provides a sound asset management method based on big data, comprising: obtaining a sound asset storage request, wherein the sound asset storage request includes an initial audio material and a target emotional style; determining the audio semantic features corresponding to the initial audio material based on the emotional expression information, theme keywords and audio structure information of the initial audio material, and screening a plurality of matching sound assets whose corresponding semantic similarities exceed a preset threshold from a sound asset library based on the audio semantic features; constructing a sound asset graph structure based on the initial audio material and each of the matching sound assets; the sound asset graph structure includes a first graph node, a plurality of second graph nodes and a plurality of second graph nodes between each of the second graph nodes and the target emotional style; The first graph nodes are connected by edges; the node features of the first graph nodes are defined by the initial audio material, the node features of each of the second graph nodes are respectively defined by the corresponding matching sound assets, and the weights of each of the edge connections are respectively defined by the semantic similarity and emotional relevance indicated by the connected graph nodes; the sound asset graph structure is processed based on a graph neural network to perform feature fusion through information transmission between nodes and edge connections, and the node features of the first graph nodes are updated; the updated node features of the first graph nodes and the target emotional style are input into a generative adversarial network to generate corresponding scene-derived sound assets, and the scene-derived sound assets are stored in the sound asset library.
[0007] In a second aspect, an embodiment of the present application provides a sound asset management system based on big data, comprising: an entry request acquisition unit for acquiring a sound asset entry request, wherein the sound asset entry request includes an initial audio material and a target emotional style; a sound asset matching unit for determining the audio semantic features corresponding to the initial audio material based on the emotional expression information, thematic keywords and audio structure information of the initial audio material, and screening a plurality of matching sound assets whose corresponding semantic similarities exceed a preset threshold from the sound asset library based on the audio semantic features; a graph structure construction unit for constructing a sound asset graph structure based on the initial audio material and each of the matching sound assets; the sound asset graph structure includes a first graph node, a plurality of second graph nodes and a first graph node in each of the second graph nodes. The node is connected to the edge of the first graph node; the node feature of the first graph node is defined by the initial audio material, the node features of each second graph node are respectively defined by the corresponding matching sound assets, and the weight of each edge connection is respectively defined by the semantic similarity and emotional relevance indicated by the connected graph nodes; a graph node updating unit is used to process the sound asset graph structure based on a graph neural network, so as to perform feature fusion through information transmission between nodes and edge connections, and update the node features of the first graph node; a sound asset warehousing unit is used to input the updated node features of the first graph node and the target emotional style into the generative adversarial network to generate corresponding scene-derived sound assets, and store the scene-derived sound assets in the sound asset library.
[0008] In a third aspect, an electronic device is provided, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the big data-based sound asset management method of any embodiment of the present application.
[0009] In a fourth aspect, an embodiment of the present application provides a storage medium on which a computer program is stored. When the program is executed by a processor, the steps of the sound asset management method based on big data of any embodiment of the present application are implemented.
[0010] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the big data-based sound asset management method of any embodiment of the present application.
[0011] The big data-based sound asset management method provided by this application can produce at least the following technical effects:
[0012] (1) By introducing advanced technologies such as big data analysis, graph neural networks, and generative adversarial networks, the semantic feature extraction and emotional style matching of sound assets are automatically performed, which significantly improves the intelligence level of sound asset management. The system can quickly screen out materials that meet specific emotional and semantic requirements from massive sound assets, realize the reorganization and application of stored sound assets, and generate highly customized audio content that meets the target emotional style.
[0013] (2) By constructing a sound asset graph structure, the initial audio material is connected with multiple matching sound assets in terms of semantic and emotional associations, and the graph neural network is used to process the graph structure. This can better explore the potential relationship between the audio material requested to be stored and the existing sound assets, realize the complex reorganization and generation of sound assets, and support the generation of richer and more personalized scene-derived sound assets.
[0014] (3) By utilizing generative adversarial networks, using emotional style as input parameters, and combining semantic features for feature fusion, we can generate sound assets that match the target emotional style. This ensures that the generated audio content is not only highly consistent with the initial material and matching material in terms of semantics, but also enhances the emotional expression ability of the stored audio, generates highly customized scene-derived sound assets, and greatly improves the application value of sound assets.
[0015] Through this technical solution, the intelligent management level and application value of sound assets have been greatly improved through advanced AI technology, so that the digital management, innovative generation and transaction flow of sound assets can be completed in a highly integrated system, providing important support for the digital protection and innovative application of national sound assets, and helping the development of the National Soundscape Port construction project. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0017] Figure 1 A flowchart showing an example of a method for managing sound assets based on big data according to an embodiment of the present application is shown;
[0018] Figure 2 A schematic diagram showing a structural connection of an example of a semantic extraction network according to an embodiment of the present application is shown;
[0019] Figure 3A schematic diagram showing a structural connection of an example of a cross-modal fusion model according to an embodiment of the present application is shown;
[0020] Figure 4 A schematic diagram showing a structural connection of an example of a generative adversarial network according to an embodiment of the present application is shown;
[0021] Figure 5 A schematic diagram showing the structural connection of an example of a feature encoding unit according to an embodiment of the present application is shown;
[0022] Figure 6 A schematic diagram showing the structural connection of an example of a feature decoding unit according to an embodiment of the present application is shown;
[0023] Figure 7 A structural block diagram of an example of a sound asset management device based on big data according to an embodiment of the present application is shown;
[0024] Figure 8 This is a schematic structural diagram of an embodiment of an electronic device of the present application. DETAILED DESCRIPTION
[0025] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0026] In the technical solutions of this application, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved shall comply with the provisions of relevant laws and regulations and shall not violate public order and good morals.
[0027] Figure 1 A flowchart of an example of a sound asset management method based on big data according to an embodiment of the present application is shown.
[0028] Regarding the execution subject of the method of the embodiment of the present application, it can be any controller or processor with computing or processing capabilities. Based on the matching mechanism of semantic similarity and emotional relevance, the most relevant audio materials can be automatically screened out from the sound asset library, reducing the time and cost of manual screening. At the same time, the construction of the graph structure and the fusion of features make the reuse of sound assets more reasonable and efficient, thereby improving the utilization efficiency of sound assets in diversified application scenarios. By using the graph neural network to update the node features in the graph structure, the semantic and emotional relationships between multiple audio materials can be fully considered, thereby providing more accurate input for the generative adversarial network. In this way, when generating sound assets in complex scenes, the emotional and semantic characteristics of the scene can be better reflected, making the stored sound assets more realistic and expressive of personalized scenes.
[0029] In some examples, it can be integrated into an electronic device or terminal through software, hardware, or a combination of software and hardware, and the type of the terminal or electronic device can be diverse, such as a mobile phone, tablet computer, or desktop computer, etc.
[0030] like Figure 1 As shown, in step S110, a sound asset storage request is obtained, and the sound asset storage request includes initial audio material and target emotional style.
[0031] In some embodiments, the execution object of the method of this embodiment is usually a server or cloud processing platform, which has the ability to interact with user terminals (such as mobile phones, computers, etc.). Specifically, communication is established with the user terminal through an application programming interface (API) to ensure secure data transmission and real-time response. In combination with business application scenarios, users can access the system platform through their terminal devices (such as smartphones, tablets, PCs, etc.). The terminal device provides a functional interface for recording or uploading audio materials and supports users to select or input the target emotional style.
[0032] For example, after recording or uploading audio material, the user can select or input the target emotional style through the application interface. These emotional styles can be diverse, such as lyrical, cheerful, sad, passionate, etc. In addition, if the user does not actively specify the target emotional style, the system can automatically match the default corresponding or all preset emotional styles based on the content of the audio material.
[0033] In step S120, the audio semantic features corresponding to the initial audio material are determined based on the emotional expression information, theme keywords and audio structure information of the initial audio material, and multiple matching sound assets whose corresponding semantic similarities exceed a preset threshold are screened from the sound asset library based on the audio semantic features.
[0034] Here, we conduct an in-depth analysis of the initial audio material to extract emotional expression information, topic keywords, and audio structure information. These three elements are then combined to generate audio semantic features, which are then used to filter multiple matching sound assets from the sound asset library. It should be understood that emotional expression information, topic keywords, and audio structure information can be obtained using various existing or potential audio content analysis technologies, and are not limited here.
[0035] In some examples of the embodiments of the present application, the emotion expression information includes the emotion expression type and emotion intensity. By using a pre-trained emotion classification model to perform emotion analysis on the initial audio material, it can be based on deep learning technology and can identify the emotion expression type in the audio, such as joy, sadness, anger, calmness, etc. Its training data includes a large number of labeled audio samples to ensure its ability to recognize different emotions. In addition, after identifying the emotion type, the emotion intensity can be further calculated by analyzing the volume, rhythm changes, pitch and other features of the audio. Emotional intensity is a quantification of emotional expression, which indicates the prominence of emotion in the audio, such as "strong sadness" or "slight joy", to achieve more personalized and accurate emotional expression.
[0036] In some examples of the embodiments of the present application, the audio structure information includes rhythmic patterns, harmonic structures and melodic trends. It should be noted that rhythmic patterns are rhythmic features in audio, such as beats, speed and regularity of rhythm; in music, harmony refers to the relationship and interaction between multiple notes sounded at the same time; melody is the sequence of notes in music, and the melodic trend describes how these notes change over time, including whether they rise, fall or remain unchanged. Specifically, by using a method that combines spectral analysis and time domain analysis, the structural features of the audio are extracted, which include the rhythmic pattern, melodic trend, harmonic structure, etc. of the audio. By analyzing these features, the system can construct a "structural fingerprint" of the audio to characterize the overall structure of the audio. On the basis of the extraction of basic structural features, hierarchical structural analysis can be further performed to decompose the audio into multiple levels (such as prelude, main song, climax, etc.). The features of each level will be extracted and stored, which will help to more accurately match sound assets with similar structures.
[0037] Regarding the process of extracting topic keywords, for example, natural language processing (NLP) technology is used to extract topic keywords from the text description of the audio material (such as lyrics, dubbing manuscripts). Specifically, through methods such as TF-IDF (term frequency–inverse document frequency) and named entity recognition (NER), keywords closely related to the audio theme can be identified. In addition, based on keyword extraction, the extracted text keywords are further associated with specific sound effects and melody fragments in the audio through cross-modal association technology. For example, an audio with an "ocean" theme may contain specific sound effects such as wave sounds and seabird calls. These sound effects are associated with keywords so that this information can be considered in subsequent processing.
[0038] Furthermore, emotional expression information, topic keywords, and audio structure information are fused to generate comprehensive audio semantic features of the initial audio material. For example, multimodal feature fusion technology is used to map different types of features (such as text and audio features) into the same semantic space, thereby forming a unified feature representation. Then, using a similarity calculation method based on semantic space, the generated audio semantic features are compared with the audio semantic features of each sound asset in the sound asset library, and the semantic similarity is calculated, thereby screening out multiple sound assets that are most similar to the initial audio material.
[0039] In step S130 , a sound asset graph structure is constructed based on the initial audio material and each matching sound asset.
[0040] Here, the sound asset graph structure includes a first graph node, multiple second graph nodes, and edge connections between each second graph node and the first graph node; the node features of the first graph node are defined by the initial audio material, the node features of each second graph node are defined by the corresponding matching sound asset, and the weight of each edge connection is defined by the semantic similarity and emotional relevance indicated by the connected graph nodes.
[0041] In some embodiments, when constructing a sound asset graph structure, each graph node represents a sound asset (or initial audio material), and the weight of the edge connection represents the strength of the relationship between the nodes, which is determined by two key factors: semantic similarity and emotional relevance. By taking semantic similarity and emotional relevance as the determining factors of edge weights, the intrinsic connection between different sound assets can be more accurately reflected. In addition, it should be noted that in the process of determining the audio semantic features, the emotional expression information of the initial audio material has been integrated, and in addition to the semantic similarity, the edge weight further integrates the emotional association analysis results, which can effectively enhance the proportion of the emotional analysis elements between the sound assets and the materials.
[0042] In some examples of the embodiments of the present application, the weight of the edge connection is calculated by the following formula:
[0043] w ij =β1·sim(F i ,F j )+β2·corr(E i ,E j ), Formula (1)
[0044]
[0045]
[0046] Where w ij represents the edge connection weight between the first graph node i and the second graph node j, β1 and β2 are weight adjustment coefficients respectively; F i and F j Represents the semantic feature vectors corresponding to i and j respectively, sim(F i ,F j ) represents the semantic similarity between i and j, ‖F i ‖ and ‖F j ‖ respectively represent F i and F j The Euclidean norm of corr(E i ,E j ) represents the emotional correlation between i and j, E i and E j Respectively represent the feature vectors of the emotional expression information corresponding to i and j; Cov(E i ,E j ) represents E i With E j The covariance between and Respectively represent E i and E j The standard deviation of .
[0047] Regarding the explanation of the above formula, the semantic similarity measure sim(·) indicates the degree of similarity between two nodes in the semantic feature space, which is expressed by cosine similarity. The higher the value of this part, the closer the semantic content of the two nodes is, and the more suitable they are for a closer connection. The emotional correlation corr(·) indicates whether the emotional expressions of the two nodes are consistent. It is measured by calculating the Pearson correlation coefficient between the emotional feature vectors. The Pearson correlation coefficient is used to measure the degree of linear correlation between two variables. The value range is between [-1,1]. For example, the closer the value is to 1, the stronger the positive correlation is, that is, the more consistent the emotional expression is. In addition, the use of dynamically adjustable weight coefficients β1 and β2 can flexibly adapt to different application scenarios, so that the graph structure can better express the diversity and complexity of sound assets.
[0048] More specifically, for each pair of candidate nodes i and j, the edge weight w between them is calculated using the above formula ij , which can represent the closeness of the edge connection between the graph nodes. In addition, in practical applications, a threshold τ can be set, and only when w ij Only when ≥τ, the edge connections between nodes are retained in the graph, which can help reduce the complexity of the graph and retain the most important connection relationships.
[0049] Through this embodiment, semantic similarity and emotional relevance are combined, and the edge connection weight can more accurately reflect the true relationship between the original audio material and the matching sound asset, so that the constructed sound asset graph structure can not only express the consistency of content, but also maintain the coordination of emotional atmosphere, thereby enhancing the overall expressiveness of the graph structure.
[0050] In step S140 , the sound asset graph structure is processed based on the graph neural network to perform feature fusion through information transfer between nodes and edges, and update the node features of the first graph node.
[0051] Here, the constructed sound asset graph structure is processed by applying a Graph Neural Network (GNN) to transfer information and fuse features between graph nodes, thereby updating the node features of the first graph node. It should be understood that GNN models can be diverse, such as Graph Attention Networks (GAT) or Graph Convolutional Networks (GCN), among others.
[0052] It should be noted that because the weights of edge connections reflect the deep semantic and emotional connections between nodes, graph neural networks can more effectively utilize this weight information when transferring and fusing features. Specifically, through the hierarchical structure of the GNN, the semantic and emotional information of the second graph node is gradually transferred to the first graph node, ultimately forming a comprehensive, updated node feature that incorporates the characteristics of the matching sound asset.
[0053] In step S150 , the updated node features of the first graph node and the target emotional style are input into a generative adversarial network to generate corresponding scene-derived sound assets, and the scene-derived sound assets are stored in a sound asset library.
[0054] Here, by inputting the node features of the updated first graph node together with the target emotional style specified by the user into the Generative Adversarial Network (GAN), the GAN generates a new scene-derived sound asset based on these inputs. The generated audio content not only maintains the semantic and structural characteristics of the original audio material, but also fully reflects the emotional style required by the user. In some embodiments, an autoregressive model can also be used to arrange and combine the generated sound clips in time sequence to ensure that the generated sound work is coherent and logical. Finally, the generated scene-derived sound asset is stored in the sound asset library for subsequent use or trading by users.
[0055] It should be noted that GAN consists of two parts: the generator and the discriminator. The generator generates derivative sound assets that meet the scene requirements by performing multi-layer processing on the input node features and the target emotional style. The discriminator is mainly used in the training and optimization stage of GAN to ensure that the sound assets generated by the generator are in line with the expected emotional style.
[0056] In some examples of the embodiments of the present application, the semantic features of the sound assets corresponding to the scene-derived sound assets are extracted based on the semantic extraction network, and the scene-derived sound assets and the corresponding sound asset semantic features are associated and stored in the sound asset library. In combination with the business scenario, when a new sound asset is entered into the sound asset library, the system first uses the semantic extraction network to analyze the sound asset, extracts its corresponding semantic features, and associates the extracted semantic features with the corresponding sound assets into the library. Thus, when the user uploads the initial audio material and requests to be stored in the library, it needs to perform semantic similarity matching. The system can first calculate the semantic features of the audio material and directly perform semantic similarity calculations with each pre-calculated and stored semantic feature. Thus, through the pre-calculation method, the response time of the semantic similarity matching is greatly shortened, so that the system can efficiently retrieve similar sound assets in the massive sound asset library. In addition, the strategy of pre-calculating and storing semantic features greatly reduces the demand for real-time computing and reduces the computing resource consumption of the system under high concurrency conditions.
[0057] Through the embodiments of the present application, GAN can generate highly customized audio content based on user needs. It is not only highly consistent with the original material and matching sound assets in terms of semantics and emotion, but also can meet the application requirements of specific scenarios, further enriching the diversity and practicality of sound assets, and greatly enhancing the application value of sound assets.
[0058] In some examples of the embodiments of the present application, audio semantic features are determined based on a semantic extraction network. Figure 2 A structural connection diagram of an example of a semantic extraction network according to an embodiment of the present application is shown.
[0059] like Figure 2 As shown, the semantic extraction network 200 includes a sentiment recognition model 210, a topic extraction model 220, a structure extraction model 230 and a cross-modal fusion model 240.
[0060] The emotion recognition model 210 is used to identify the feature vector of the emotion expression information corresponding to the initial audio material, the topic extraction model 220 is used to extract the feature vector of the topic keyword corresponding to the initial audio material, the structure extraction model 230 is used to extract the feature vector of the audio structure information corresponding to the initial audio material, and the cross-modal fusion model 240 is used to perform cross-modal feature fusion on the feature vector of the emotion expression information, the feature vector of the topic keyword and the feature vector of the audio structure information to obtain audio semantic features.
[0061] In some embodiments, the topic extraction model 220 extracts representative keywords from audio-related text (e.g., text obtained through speech recognition, lyrics, dubbing scripts, etc.), and converts these keywords into topic feature vectors through a word vector model (e.g., Word2Vec), which can represent the main semantic topics of the audio content.
[0062] In some embodiments, the structural features of the audio, such as spectrum, rhythm, harmony, etc., are analyzed by the structure extraction model 230 to generate a structural feature vector, which describes the overall structure and organization of the audio, such as the speed of the rhythm, the direction of the melody, etc. Exemplarily, the structure extraction model 230 can adopt a multi-level convolutional neural network, and each layer of convolution kernel has a different receptive field to capture features of different scales. In the hierarchical structure, through the first layer of convolution, a smaller convolution kernel (such as 3×3) is used to extract local rhythm pattern features in the audio signal; through the second layer of convolution, a medium-sized convolution kernel (such as 5×5) is used to capture harmonic structure features, including chord changes and pitch combinations; through the third layer of convolution, a larger convolution kernel (such as 7×7) is used to identify the melody direction features and extract the changing trend of the note sequence. Thus, through the multi-level convolution structure, it is possible to capture structural information of different scales and levels in the audio signal, and perform a full-scale analysis from rhythm to melody, making the structural feature extraction more comprehensive and accurate.
[0063] In some examples of the embodiments of the present application, the emotion recognition model 210 adopts a self-supervised learning model, which is optimized through multiple self-supervised tasks in the pre-training stage. The multiple self-supervised tasks include time suppression prediction tasks and spectrum filling tasks, thereby helping the model learn more robust and general feature representations.
[0064] More specifically, in the time suppression prediction task, some time domain features in the time domain audio signal of the sample are randomly masked, and the emotion recognition model predicts the masked time domain features, so that the model can better understand the dependencies in the time series. In the spectrum filling task, some frequency domain features in the spectrum graph of the sample are randomly suppressed, and the emotion recognition model predicts the masked frequency domain features, so that the model is better at processing frequency domain information and understanding the frequency structure of the audio signal. Furthermore, after self-supervised learning pre-training and feature fusion, the emotion recognition model enters the final emotion recognition stage, including the classification of emotion types and the prediction of emotion intensity. In this way, before the model formally performs the emotion recognition task, it is first pre-trained using unlabeled audio data through a self-supervised task, which can help the model learn a general feature representation and improve feature extraction capabilities.
[0065] Therefore, by performing self-supervised task training on unlabeled data, the model can learn to extract key time-domain and frequency-domain features from audio signals, allowing the model to be effectively trained without a large amount of labeled data, reducing the cost of sample labeling and training. In addition, through self-supervised pre-training, the model is in a good state at the beginning of formal emotion classification training, and the model has mastered a certain level of feature extraction capabilities, which can effectively accelerate the convergence of emotion recognition task training and improve the final recognition accuracy.
[0066] Figure 3 A structural connection diagram of an example of a cross-modal fusion model according to an embodiment of the present application is shown.
[0067] like Figure 3 As shown, the cross-modal fusion model 240 includes a cascaded input layer 310, a feature alignment layer 320, a self-attention fusion layer 330 and an MLP (Multilayer Perceptron) layer 340.
[0068] More specifically, the input layer 310 is used to receive feature vectors of each modality. Specifically, the input layer 310 is connected to the emotion recognition model, the topic extraction model, and the structure extraction model respectively to receive feature vectors corresponding to emotion expression information, topic keywords, and audio structure information.
[0069] The feature alignment layer 320 is used to map the feature vectors of each modality into a unified feature space by using a fully connected layer. Since the feature vector dimensions of different modalities may be different, the model first performs feature alignment to map all feature vectors into a unified feature space. To this end, a fully connected layer is used to perform a linear transformation:
[0070] F′ emotion =W emotion F emotion +b emotion , Formula (4)
[0071] F′ theme =W theme F theme +b theme , Formula (5)
[0072] F′ structure =W structure F structure +b structure , Formula (6)
[0073] Where, F emotion is the feature vector of emotional expression information, with dimension d e ; F theme The feature vector representing the topic keywords, with dimension d t; F structure The feature vector representing the audio structure information has a dimension of d s ;W emotion and b emotion They represent the weight matrix and bias term of the fully connected layer used to process the feature vector of emotional expression information; W emotion The dimension is d f ×d e , b emotion The dimension is d f , where d f is the dimension of the unified feature space; W theme and b theme Respectively represent the weight matrix and bias term of the fully connected layer used to process the feature vector of the topic keyword; W theme The dimension is d f ×d t , b theme The dimension is d f ;W structure and b structure Represent the weight matrix and bias term of the fully connected layer for processing the feature vector of audio structure information; W structure The dimension is d f ×d s , b structure The dimension is d f ; F′ emotion , F′ theme and F′ structure Represents F emotion 、F theme and F structure After feature alignment, the dimension of the feature vector is unified to d f .
[0074] Through the feature alignment layer, the feature representations of different modalities are standardized to ensure that they are in the same dimensional feature space, laying the foundation for subsequent feature fusion, so that features of different modalities can be effectively compared and integrated.
[0075] The self-attention fusion layer 330 is used to calculate the similarity between different modal features through the self-attention mechanism, and dynamically adjust the attention weight of each modal feature according to the similarity, and perform weighted summation of each modal feature to determine the corresponding fused semantic feature. Here, the self-attention mechanism is introduced to capture the mutual relationship and importance between different modal features. For each input feature vector, its correlation with other modal feature vectors is calculated, thereby dynamically adjusting the weight of each modal feature, and applying the attention weight to the modal feature, and fusion is performed by weighted summation:
[0076]
[0077] Where F′ m and F′ p Respectively represent the mth modal eigenvector and the pth modal eigenvector after alignment; sim(F′ m ,F′ p ) represents F′ m and F′ p similarity between is a normalization factor, which represents the comprehensive similarity of all modal eigenvectors satisfying s≠m; α mp In the self-attention mechanism, F represents the attention weight of the mth modal feature vector to the pth modal feature vector; fused Represents the fusion modal features, dimension d f .
[0078] Through the self-attention fusion layer, the importance of each modal feature to the final semantic expression is dynamically determined, and feature fusion is performed based on the importance, so that the model can adaptively adjust the influence of each modal feature, ensuring that the final fused feature vector can most accurately express the comprehensive semantic characteristics of the initial audio material.
[0079] Here, the introduction of the self-attention mechanism enables the model to adaptively assign weights to each modality, ensuring that the most important modal features are prioritized in each specific audio scenario. For example, in audio material with strong emotional expression, the weight of emotional features may be dynamically increased, while in structurally complex music clips, audio structural information may dominate, allowing the final semantic feature representation to better reflect the core characteristics of the audio material.
[0080] The MLP layer 340 is used to perform linear transformation and nonlinear activation processing on the fused modal features layer by layer to determine the corresponding audio semantic features. To further extract high-order semantic features, the fused modal features are input to the MLP layer, which consists of multiple fully connected layers and nonlinear activation functions (such as ReLU) to extract deep features layer by layer.
[0081] F g+1 =RELU(W g F g +b g ), Formula (10)
[0082] F output =RELU(W L F L +b L ), Formula (11)
[0083] Where RELU represents the RELU nonlinear activation function; W g and bg Represent the weight matrix and bias term of the g-th fully connected layer, F g and F g+1 They represent the input features and output features of the g-th layer respectively; L is the total number of MLP layers, F output Represents the audio semantic features; W L 、b L and F L They represent the weight matrix, bias term and input features of the Lth fully connected layer respectively.
[0084] Through the MLP layer, deep learning is used to further explore the complex semantic relationships in the fusion features, which can capture the nonlinear relationship between features and make the output comprehensive semantic features more expressive and accurate.
[0085] Through the embodiments of the present application, a combination of MLP layers and self-attention mechanisms is adopted, so that the model can capture the complex relationships and interactions between them when fusing features of different modalities. Through deep fusion, it can not only integrate emotional, thematic and structural features, but also enhance the understanding and expression capabilities of complex semantic relationships through the extraction of high-order features.
[0086] In some examples of the embodiments of the present application, the graph neural network can adopt a temporal attention graph neural network.
[0087] Specifically, for each edge connection, the time dynamic weight w is calculated in combination with the time information ij (t):
[0088] w ij (t) = w ij ·exp(-λ·Δt), Formula (12)
[0089] Where w ij (t) represents the time-dynamic edge weight calculated in combination with time information; λ is the time attenuation coefficient, which is used to control the degree of influence of time on the weight; Δt is the difference between the audio recording time of the initial audio material indicated by the first graph node i and the storage time of the matching sound asset indicated by the second graph node j.
[0090] Here, by introducing the time dynamic weight w ij(t), the model can dynamically adjust the weights of nodes and edges to reflect the changing characteristics of audio materials and sound assets at different time points, allowing the model to better capture the characteristics of audio materials that change over time and ensure that the generated sound assets conform to the current time state and trends. In addition, the time decay mechanism ensures that the influence of older sound assets on the current feature update gradually weakens, while assets closer to the current time contribute more to the features. This achieves time-sensitive feature weighting calculation, helps generate sound assets that better meet current application needs, and improves the accuracy of system responses.
[0091] Calculate the attention weight α between nodes based on the graph attention mechanism ij (t):
[0092] e ij (t) = LeakyReLU(a T [WF i (t)‖WF j (t)]), Formula (13)
[0093]
[0094] Where, F i (t) and F j (t) are the feature representations of i and j at time t, W is the linear transformation matrix, a T is the transpose of the attention vector a, e ij (t) represents the attention score of i and j at time t, LeakyReLU represents the LeakyReLU activation function, ‖ represents the vector connection operation; N(i) represents the set of neighbor nodes of i, α ij (t) represents the normalized attention weight of i to j at time t;
[0095] is a normalization factor, which represents the sum of the attention scores of all neighbor nodes k of node i at time t after exponential operation.
[0096] By introducing the attention mechanism, the model can adaptively adjust the influence of neighboring nodes on target node feature updates based on the feature similarity and importance between nodes. This allows it to prioritize sound assets that are more semantically and emotionally relevant to the initial audio material during feature fusion, generating more accurate node feature representations. This effectively suppresses interference from noise and irrelevant nodes, making the feature update process more stable and accurate, ensuring the model maintains efficient feature extraction and fusion capabilities when processing complex and diverse audio data.
[0097] Combining the temporal dynamic weight and the attention weight, the neighbor node features are weighted summed to update the feature representation of the first graph node i:
[0098]
[0099] Where, represents the updated feature representation of the first graph node i at time t.
[0100] Through the embodiments of the present application, the temporal dynamic information is combined with the attention mechanism, which can simultaneously integrate multi-dimensional information such as time, semantics and emotion in the feature representation. By introducing temporal dynamic weights, outdated or less relevant sound assets can be effectively filtered out, reducing unnecessary computational burden. In addition, the attention mechanism further optimizes the allocation of computing resources, allowing the system to maintain high accuracy and stability while efficiently processing large amounts of data. As a result, key features in audio materials can be extracted and integrated more intelligently, making the management and generation process of sound assets more intelligent, especially in terms of emotional expression and time sensitivity, so that the final generated sound assets can better match personalized emotional style requirements.
[0101] In some examples of the embodiments of the present application, the generative adversarial network adopts a variational auto-encoder (VAE) generative adversarial network based on attention enhancement. Here, VAE-GAN combines the advantages of variational autoencoder (VAE) and generative adversarial network (GAN), using the VAE part to learn the potential representation of the input data and generate new samples; while the GAN part is used to improve the authenticity of the generated samples, making them indistinguishable from real samples. In this way, the diversity and authenticity of the generated content can be guaranteed at the same time, and by controlling the latent space, audio assets that conform to a specific emotional style can be generated.
[0102] Figure 4 A schematic diagram of the structural connection of an example of a generative adversarial network according to an embodiment of the present application is shown.
[0103] The generator 400 of the generative adversarial network includes a cascaded input unit 410, an encoder 420, and a decoder 430. The encoder 420 includes a plurality of cascaded feature encoding units (4211, 4212...421m) and a latent representation generation unit 422, and the decoder 430 includes a plurality of cascaded feature decoding units (4311, 4312...431n) and an audio generation unit 432.
[0104] The input unit 410 is used to fuse the updated node features of the first graph node and the feature vector of the target emotional style:
[0105]
[0106] Where, E target The feature vector representing the target emotional style, F input represents the comprehensive input feature vector.
[0107] In this way, by introducing the emotion control mechanism in the VAE part, the generator can generate audio assets that conform to a specific emotional style to meet the personalized needs of users.
[0108] Figure 5 FIG. 1 shows a schematic diagram of a structural connection of an example of a feature coding unit according to an embodiment of the present application. Figure 5 As shown, each feature encoding unit 500 includes a cascaded first convolution layer 510, an encoding multi-head attention layer 520 and a second convolution layer 530.
[0109] The first convolutional layer 510 is used to extract local patterns of input features through a sliding window:
[0110]
[0111] Where, is the first convolutional layer operation of the lth feature encoding unit, Represents the output feature map of the first convolutional layer, F input,l-1 Indicates F input Or the output feature map of the l-1th feature encoding unit.
[0112] The encoding multi-head attention layer 520 is used to process the feature map through the multi-head attention mechanism:
[0113]
[0114] In the formula, MultiHead means that the multi-head attention mechanism learns different feature relationships through multiple parallel attention heads. The query vector, key vector, and value vector are used as the attention mechanism at the same time; Represents the output feature map of the encoded multi-head attention layer of the l-th feature encoding unit.
[0115] The second convolutional layer 530 is used to compress the spatial dimension of the features to extract higher-level patterns:
[0116]
[0117] Where, represents the second convolutional layer operation of l feature encoding units, Represents the output feature map of the second convolutional layer.
[0118] By combining multi-layer convolution with a multi-head attention mechanism in the encoder, the model can extract and fuse features of audio signals at different levels. This ensures that the key information of the audio signal can be effectively extracted through layer-by-layer extraction and compression.
[0119] The potential representation generation unit 422 is used to generate a potential representation according to the output feature map of the last feature encoding unit:
[0120]
[0121] z Q =μ+σ VAE ·ò, formula (22)
[0122] Where μ and σ VAE Represent the mean vector and standard deviation vector of the latent space, which are respectively generated by the corresponding convolution operation Conv μ and Conv σ Calculated; z Q represents the latent space representation, represents random noise from a standard normal distribution; Represents the output feature map of the Qth feature coding unit, where Q is the total number of feature coding units.
[0123] Here, the encoder reduces the dimensionality of the input features through multiple convolutional layers, gradually compressing the dimensions of the feature map and extracting the latent representation. In VAE, the representation of the latent space is determined by the mean vector and standard deviation vector of the latent space.
[0124] This embodiment, based on the structure of a variational autoencoder, introduces randomness into the latent space of the generator, ensuring diversity in the generated audio signals. This allows the model to explore more regions of the latent space when generating audio and generate audio samples of different styles. Furthermore, this enhances the model's robustness to input noise and uncertainty, ensuring that the generated audio signals maintain high quality across a variety of input scenarios.
[0125] Figure 6 FIG. 1 shows a schematic diagram of a structural connection of an example of a feature decoding unit according to an embodiment of the present application. Figure 6 As shown, each feature decoding unit 600 includes a cascaded first transposed convolution layer 610, a decoding multi-head attention layer 620 and a second transposed convolution layer 630.
[0126] The decoder gradually restores the audio signal through serial decoding units. Combined with the multi-head attention mechanism, it can continuously focus on and retain important semantic and emotional information during the decoding process, ensuring that important features are not lost during the decoding process. The final generated audio signal can achieve a higher level in quality and consistency.
[0127] The first transposed convolutional layer 610 is used to gradually restore the spatial dimensions of the feature map through deconvolution operations:
[0128]
[0129] Where, represents the output feature map of the first transposed convolutional layer of the nth feature decoding unit, represents the first deconvolution operation of the nth feature decoding unit, z Q,n-1 represents z Q Or the output feature map of the n-1th feature decoding unit.
[0130] The decoding multi-head attention layer 620 is used to weightedly summarize different feature patterns:
[0131]
[0132] Where, As the query vector for decoding the multi-head attention mechanism, E target As the key vector and value vector of the decoding multi-head attention mechanism, Represents the output feature map of the decoding multi-head attention layer.
[0133] The second transposed convolution layer 630 is used to restore the detailed information in the audio signal through multi-layer deconvolution operations:
[0134]
[0135] Where, represents the second deconvolution operation, Represents the output feature map of the second transposed convolutional layer of the nth feature encoding unit.
[0136] Here, the decoder reconstructs the latent space representation into audio data through the deconvolution layer (or transposed convolution layer), and in the decoding process, the target emotional style vector E is converted into target As conditional information, it is fed into the decoder together with the latent space representation to ensure that the generated audio is emotionally consistent with user needs.
[0137] Furthermore, the multi-head attention mechanism ensures that the generated audio signal is not only structurally sound but also accurately expresses the target emotion by focusing on the target emotion vector multiple times during the decoding process. This allows the generator to achieve greater consistency in emotional expression, ensuring that the generated audio signal closely matches the input emotion. For example, if the target emotion is "happiness," the generated audio signal will reflect this emotional characteristic in terms of rhythm, pitch, and other aspects.
[0138] The audio generation unit 432 is used to convert the output feature map of the last feature decoding unit into a scene-derived sound asset:
[0139]
[0140] Where A recon Represents the generated scene-derived sound asset, TransConv (N+1) represents the deconvolution operation of the audio generation unit, Represents the output feature map of the Nth feature decoding unit, where N is the total number of feature decoding units.
[0141] By introducing the multi-head attention mechanism, the generator can capture the complex relationship between input features in the encoder and decoder, and can dynamically focus on the key features in the audio signal, thereby improving the details and overall sound quality of the generated audio signal. The generator not only performs well in feature extraction and information compression, but can also effectively retain information during the decoding process, making the generated audio more layered and clear, and can more accurately restore the semantic features and emotional expression of the audio signal.
[0142] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of combined actions, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application. In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0143] Figure 7 A structural block diagram of an example of a sound asset management device based on big data according to an embodiment of the present application is shown.
[0144] like Figure 7 As shown, the sound asset management device 700 based on big data includes a storage request acquisition unit 710 , a sound asset matching unit 720 , a graph structure construction unit 730 , a graph node updating unit 740 and a sound asset storage unit 750 .
[0145] The storage request acquisition unit 710 is used to acquire a sound asset storage request, where the sound asset storage request includes an initial audio material and a target emotional style.
[0146] The sound asset matching unit 720 is used to determine the audio semantic features corresponding to the initial audio material based on the emotional expression information, theme keywords and audio structure information of the initial audio material, and to screen multiple matching sound assets whose corresponding semantic similarities exceed a preset threshold from the sound asset library based on the audio semantic features.
[0147] The graph structure construction unit 730 is used to construct a sound asset graph structure based on the initial audio material and each of the matching sound assets; the sound asset graph structure includes a first graph node, multiple second graph nodes, and edge connections between each of the second graph nodes and the first graph node; the node features of the first graph node are defined by the initial audio material, the node features of each of the second graph nodes are defined by the corresponding matching sound assets, and the weights of each of the edge connections are defined by the semantic similarity and emotional relevance indicated by the connected graph nodes.
[0148] The graph node updating unit 740 is configured to process the sound asset graph structure based on a graph neural network, perform feature fusion through information transfer between nodes and edges, and update node features of the first graph nodes.
[0149] The sound asset storage unit 750 is used to input the updated node features of the first graph node and the target emotional style into the generative adversarial network to generate corresponding scene-derived sound assets, and store the scene-derived sound assets in the sound asset library.
[0150] In some embodiments, an embodiment of the present application provides a non-volatile computer-readable storage medium, which stores one or more programs including execution instructions, and the execution instructions can be read and executed by electronic devices (including but not limited to computers, servers, or network devices, etc.) to execute the steps of any of the above-mentioned big data-based sound asset management methods of the present application.
[0151] In some embodiments, the embodiments of the present application also provide a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer performs any step of the above-mentioned big data-based sound asset management method.
[0152] In some embodiments, an embodiment of the present application also provides an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the big data-based sound asset management method.
[0153] Figure 8 This is a hardware structure diagram of an electronic device that implements a sound asset management method based on big data, as provided in another embodiment of the present application. Figure 8 As shown, the device includes:
[0154] One or more processors 810 and memory 820, Figure 8 A processor 810 is taken as an example.
[0155] The device for executing the big data-based sound asset management method may further include: an input device 830 and an output device 840 .
[0156] The processor 810, the memory 820, the input device 830 and the output device 840 may be connected via a bus or other means. Figure 8 The bus connection is taken as an example.
[0157] Memory 820, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the big data-based sound asset management method in the embodiments of the present application. Processor 810 executes the non-volatile software programs, instructions, and modules stored in memory 820 to execute various server functional applications and data processing, thereby implementing the big data-based sound asset management method in the aforementioned method embodiment.
[0158] The memory 820 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device, etc. In addition, the memory 820 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 820 may optionally include a memory remotely located relative to the processor 810, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0159] The input device 830 can receive input digital or character information and generate signals related to user settings and function control of the electronic device. The output device 840 can include a display device such as a display screen.
[0160] The one or more modules are stored in the memory 820 and, when executed by the one or more processors 810 , perform the sound asset management method based on big data in any of the above method embodiments.
[0161] The above-mentioned product can execute the method provided in the embodiment of this application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the method provided in the embodiment of this application.
[0162] The electronic devices of the embodiments of the present application exist in various forms, including but not limited to:
[0163] (1) Mobile communication devices: These devices are characterized by their mobile communication capabilities and their primary purpose is to provide voice and data communications. These terminals include smartphones, multimedia phones, feature phones, and low-end phones.
[0164] (2) Ultra-mobile personal computer devices: These devices fall under the category of personal computers and have computing and processing capabilities, and generally also have mobile Internet access. These terminals include PDAs, MIDs, and UMPCs.
[0165] (3) Portable entertainment devices: These devices can display and play multimedia content. They include audio and video players, handheld game consoles, e-books, smart toys, and portable car navigation devices.
[0166] (4) Other onboard electronic devices with data interaction functions, such as onboard computer devices installed in vehicles.
[0167] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0168] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, or of course, by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiment.
[0169] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A sound asset management method based on big data, comprising: Obtaining a sound asset storage request, wherein the sound asset storage request includes an initial audio material and a target emotional style; Determining audio semantic features corresponding to the initial audio material based on the emotional expression information, theme keywords, and audio structure information of the initial audio material, and screening, from a sound asset library, a plurality of matching sound assets whose corresponding semantic similarities exceed a preset threshold based on the audio semantic features; Building a sound asset graph structure based on the initial audio material and each of the matching sound assets; The sound asset graph structure includes a first graph node, a plurality of second graph nodes, and an edge connection between each of the second graph nodes and the first graph node; the node features of the first graph node are defined by the original audio material, the node features of each of the second graph nodes are defined by the corresponding matched sound asset, and the weight of each of the edge connections is defined by the semantic similarity and emotional relevance indicated by the connected graph nodes; processing the sound asset graph structure based on a graph neural network to perform feature fusion through information transfer between nodes and edges, and updating node features of the first graph nodes; Inputting the updated node features of the first graph nodes and the target emotional style into a generative adversarial network to generate corresponding scene-derived sound assets, and storing the scene-derived sound assets in the sound asset library; The scene-derived sound asset is a sound asset generated by reorganizing and applying the sound asset library to derive the initial audio material according to the target emotional style; The audio semantic features are determined based on a semantic extraction network; the semantic extraction network includes an emotion recognition model, a topic extraction model, a structure extraction model and a cross-modal fusion model; The weight of the edge connection is calculated by the following formula: w ij =β1·sim(F i ,F j )+β2·corr(E i ,E j ), Where w ij represents the edge connection weight between the first graph node i and the second graph node j, β1 and β2 are weight adjustment coefficients respectively; F i and F j Represents the semantic feature vectors corresponding to i and j respectively, sim(F i ,F j ) represents the semantic similarity between i and j, ‖F i ‖ and ‖F j ‖ respectively represent F i and F j The Euclidean norm of corr(E i ,E j ) represents the emotional correlation between i and j, E i and E j Respectively represent the feature vectors of the emotional expression information corresponding to i and j; Cov(E i ,E j ) represents E i With E j The covariance between and Respectively represent E i and E j The standard deviation of For each edge connection, the time dynamic weight w is calculated by combining the time information ij (t): w ij (t)=w ij ·exp(-λ·△t), Where w ij (t) represents the time dynamic edge weight calculated by combining time information; λ is the time decay coefficient, which is used to control the influence of time on the weight; Δt is the difference between the audio recording time of the initial audio material indicated by the first graph node i and the storage time of the matching sound asset indicated by the second graph node j.
2. The method according to claim 1, wherein The emotion recognition model is used to identify the feature vector of the emotion expression information corresponding to the initial audio material, the topic extraction model is used to extract the feature vector of the topic keyword corresponding to the initial audio material, the structure extraction model is used to extract the feature vector of the audio structure information corresponding to the initial audio material, and the cross-modal fusion model is used to perform cross-modal feature fusion on the feature vector of the emotion expression information, the feature vector of the topic keyword, and the feature vector of the audio structure information to obtain audio semantic features; The cross-modal fusion model includes a cascaded input layer, a feature alignment layer, a self-attention fusion layer, and an MLP layer; The input layer is used to receive each modal feature vector; The feature alignment layer is used to map the feature vectors of each modality into a unified feature space by using a fully connected layer: F′ emotion =W emotion F emotion +b emotion , F′ theme =W theme F theme +b theme , F′ structure =W structure F structure +b structure , Where, F emotion is the feature vector of emotional expression information, with dimension d e ; F theme The feature vector representing the topic keywords, with dimension d t ; F structure The feature vector representing the audio structure information has a dimension of d s ;W emotion and b emotion They represent the weight matrix and bias term of the fully connected layer used to process the feature vector of emotional expression information; W emotion The dimension is d f ×d e , b emotion The dimension is d f , where d f is the dimension of the unified feature space; W theme and b theme Respectively represent the weight matrix and bias term of the fully connected layer used to process the feature vector of the topic keyword; W theme The dimension is d f ×d t , b theme The dimension is d f ;W structure and b structure Represent the weight matrix and bias term of the fully connected layer for processing the feature vector of audio structure information; W structure The dimension is d f ×d s , b structure The dimension is d f ; F′ emotion , F′ theme and F′ structure Represents F emotion 、F theme and F structure After feature alignment, the dimension of the feature vector is unified to d f ; The self-attention fusion layer is used to calculate the similarity between different modal features through the self-attention mechanism, and dynamically adjust the attention weight of each modal feature according to the similarity, and perform weighted summation of each modal feature to determine the corresponding fused semantic feature: Where F′ m and F′ p Respectively represent the mth modal eigenvector and the pth modal eigenvector after alignment; sim(F′ m ,F′ p ) represents F′ m and F′ p similarity between is a normalization factor, which represents the comprehensive similarity of all modal eigenvectors satisfying s≠m; α mp In the self-attention mechanism, F represents the attention weight of the mth modal feature vector to the pth modal feature vector; fused Represents the fusion modal features, dimension d f ; The MLP layer is used to perform linear transformation and nonlinear activation processing on the fused modal features layer by layer to determine the corresponding audio semantic features: F g+1 =RELU(W g F g +b g ), F output =RELU(W L F L +b L ), Where RELU represents the RELU nonlinear activation function; W g and b g Represent the weight matrix and bias term of the g-th fully connected layer, F g and F g+1 Represent the input features and output features of the g-th layer respectively; L is the total number of MLP layers, F output Represents audio semantic features; W L 、b L and F L They represent the weight matrix, bias term and input features of the Lth fully connected layer respectively.
3. The method according to claim 1, wherein Storing the scene-derived sound asset in the sound asset library includes: The sound asset semantic features corresponding to the scene-derived sound asset are extracted based on the semantic extraction network, and the scene-derived sound asset and the corresponding sound asset semantic features are associated and stored in the sound asset library.
4. The method according to claim 1, wherein The emotion recognition model adopts a self-supervised learning model, which is optimized through multiple self-supervised tasks in the pre-training stage. The multiple self-supervised tasks include a time suppression prediction task and a spectrum filling task. In the time suppression prediction task, some time domain features in the time domain audio signal of the sample are randomly masked, and the emotion recognition model predicts the time domain features of the masked part. In the spectrum filling task, some frequency domain features in the spectrum graph of the sample are randomly suppressed, and the emotion recognition model predicts the frequency domain features of the masked part.
5. The method according to any one of claims 1 to 4, wherein The emotion expression information includes the emotion expression type and emotion intensity; the audio structure information includes rhythm pattern, harmony structure and melody direction.
6. The method according to claim 1, wherein The graph neural network adopts a temporal attention graph neural network and updates the node features of the first graph node in the following way: Calculate the attention weight α between nodes based on the graph attention mechanism ij (t): e ij (t)=LeakyReLU(a T [WF i (t)‖WF j (t)]), Where, F i (t) and F j (t) are the feature representations of i and j at time t, W is the linear transformation matrix, a T is the transpose of the attention vector a, e ij (t) represents the attention score of i and j at time t, LeakyReLU represents the LeakyReLU activation function, and ‖ represents the vector connection operation; represents the set of neighbor nodes of i, α ij (t) represents the normalized attention weight of i to j at time t; is a normalization factor, which represents the sum of the attention scores of all neighbor nodes k of node i at time t after exponential operation; Combining the temporal dynamic weight and the attention weight, the neighbor node features are weighted summed to update the feature representation of the first graph node i: Where, represents the updated feature representation of the first graph node i at time t.
7. The method according to claim 6, wherein: The generative adversarial network adopts an attention-enhanced variational autoencoder generative adversarial network, wherein the generator of the generative adversarial network comprises a cascaded input unit, an encoder, and a decoder; the encoder comprises a cascaded plurality of feature encoding units and a potential representation generation unit, and the decoder comprises a cascaded plurality of feature decoding units and an audio generation unit; The input unit is used to fuse the updated node features of the first graph node and the feature vector of the target emotional style: Where, E target The feature vector representing the target emotional style, F input represents the comprehensive input feature vector; Each of the feature encoding units comprises a cascaded first convolutional layer, an encoding multi-head attention layer, and a second convolutional layer; The first convolutional layer is used to extract local patterns of input features through a sliding window: Where, is the first convolutional layer operation of the lth feature encoding unit, Represents the output feature map of the first convolutional layer, F input,l-1 Indicates F input Or the output feature map of the l-1th feature encoding unit; The encoding multi-head attention layer is used to process the feature map through the multi-head attention mechanism: In the formula, MultiHead means that the multi-head attention mechanism learns different feature relationships through multiple parallel attention heads. The query vector, key vector, and value vector are used as the attention mechanism at the same time; Represents the output feature map of the encoded multi-head attention layer of the l-th feature encoding unit; The second convolutional layer is used to compress the spatial dimension of the features to extract higher-level patterns: Where, represents the second convolutional layer operation of l feature encoding units, Represents the output feature map of the second convolutional layer; The potential representation generation unit is used to generate a potential representation according to the output feature map of the last feature encoding unit: z Q =μ+σ VAE ·∈, Where μ and σ VAE Represent the mean vector and standard deviation vector of the latent space, which are respectively generated by the corresponding convolution operation Conv μ and Conv σ Calculated; z Q represents the latent space representation, ∈ represents the random noise of standard normal distribution; represents the output feature map of the Qth feature coding unit, where Q is the total number of feature coding units; Each of the feature decoding units comprises a cascaded first transposed convolutional layer, a decoding multi-head attention layer, and a second transposed convolutional layer; The first transposed convolutional layer is used to gradually restore the spatial dimensions of the feature map through deconvolution operations: Where, represents the output feature map of the first transposed convolutional layer of the nth feature decoding unit, represents the first deconvolution operation of the nth feature decoding unit, z Q,n-1 represents z Q Or the output feature map of the n-1th feature decoding unit; The decoding multi-head attention layer is used to weightedly summarize different feature modes: Where, As the query vector for decoding the multi-head attention mechanism, E target As the key vector and value vector of the decoding multi-head attention mechanism, Represents the output feature map of the decoding multi-head attention layer; The second transposed convolution layer is used to restore the detailed information in the audio signal through multi-layer deconvolution operations: Where, represents the second deconvolution operation, Represents the output feature map of the second transposed convolutional layer of the nth feature encoding unit; The audio generation unit is used to convert the output feature map of the last feature decoding unit into a scene-derived sound asset: Where A recon Represents the generated scene-derived sound asset, TransConv (N+1) represents the deconvolution operation of the audio generation unit, Represents the output feature map of the Nth feature decoding unit, where N is the total number of feature decoding units.
8. A sound asset management system based on big data, for implementing the method according to any one of claims 1 to 7, the system comprising: An entry request acquiring unit, configured to acquire a sound asset entry request, wherein the sound asset entry request includes an initial audio material and a target emotional style; a sound asset matching unit, configured to determine audio semantic features corresponding to the initial audio material based on the emotional expression information, theme keywords, and audio structure information of the initial audio material, and to select, from a sound asset library, a plurality of matching sound assets whose corresponding semantic similarities exceed a preset threshold based on the audio semantic features; A graph structure building unit, configured to build a sound asset graph structure based on the initial audio material and each of the matching sound assets; The sound asset graph structure includes a first graph node, a plurality of second graph nodes, and an edge connection between each of the second graph nodes and the first graph node; the node features of the first graph node are defined by the original audio material, the node features of each of the second graph nodes are defined by the corresponding matched sound asset, and the weight of each of the edge connections is defined by the semantic similarity and emotional relevance indicated by the connected graph nodes; a graph node updating unit, configured to process the sound asset graph structure based on a graph neural network, perform feature fusion through information transfer between nodes and edges, and update node features of the first graph nodes; The sound asset storage unit is used to input the updated node features of the first graph node and the target emotional style into the generative adversarial network to generate corresponding scene-derived sound assets, and store the scene-derived sound assets in the sound asset library.
9. A storage medium having a computer program stored thereon, wherein: When the program is executed by a processor, the sound asset management method based on big data described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Voice reminding method and device, electronic equipment and storage medium
CN117912442A
Online music activity system based on cloud computing and artificial intelligence
CN118132797A