Audio analysis method and device based on feature fusion, equipment and medium
Through feature fusion and knowledge graph analysis, the problems of insufficient multimodal information fusion and semantic understanding of audio analysis in the prior art are solved, the accuracy and interpretability of the analysis results are improved, and the correlation with industry knowledge is enhanced.
Patent Information
- Application Number
- CN202510277366.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-24
AI Technical Summary
The existing audio analytical technology has shortcomings in multimodal information fusion, semantic understanding and industry knowledge correlation, resulting in the limitation of the accuracy, interpretability and practicality of the analytical results.
The audio analysis method based on feature fusion is adopted. By obtaining the audio and associated text in the target field, the music feature vector and text semantic feature vector are extracted, and the knowledge graph is input for semantic analysis after the fusion is used to generate the analysis results.
It realizes multimodal understanding of audio data, improves the accuracy and interpretability of analysis results, and enhances the correlation between analysis results and industry knowledge, thereby improving the applicability of audio analysis technology.
Smart Images

Figure CN120197624A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to an audio parsing method, apparatus, device and storage medium based on feature fusion. Background Art
[0002] In the fields of culture and art, healthcare, and finance, the parsing and understanding of audio data have become key technologies. However, existing audio parsing technologies still have many deficiencies, especially in aspects such as multi-modal information fusion, semantic understanding, and industry knowledge association, resulting in limitations in the accuracy, interpretability, and practicality of parsing results.
[0003] In the field of culture and art, taking Buddhism as an example, the current parsing of Buddhist music mainly focuses on the analysis of basic music elements such as melody and rhythm, lacking the exploration of the deeper Buddhist cultural connotations contained in the music. For example, it can only identify the rhythm changes of chanting music, but it is difficult to understand the meanings such as Buddhist cultivation stages and mood changes expressed behind the rhythm. In addition, existing audio parsing technologies have not established effective connections between Buddhist music and cultural aspects such as Buddhist doctrines, rituals, and history. For example, when parsing Buddhist ceremony music, it is difficult to explain the specific role of the music in the ceremony process and its connection with the Buddhist thoughts conveyed in the ceremony. The parsing results often use highly specialized music theory terms, making it difficult for the general public and non-music professional Buddhist enthusiasts to understand, thus affecting the dissemination and learning of Buddhist music. In addition, Buddhist music is usually accompanied by various modal information such as scripture chanting and ritual images, but existing technologies have not effectively integrated this information. For example, when parsing Buddhist chanting music, it has not been analyzed in combination with the content of the chanted scriptures and the ritual scene, resulting in one-sided music parsing results and being unable to fully display the cultural value and artistic charm of Buddhist music.
[0004] In the field of healthcare, existing audio parsing technologies mainly focus on basic speech recognition and emotion analysis, and it is difficult to accurately understand the speech interaction information between doctors and patients. For example, in an intelligent medical consultation system, a doctor's speech diagnosis usually contains complex information such as professional terms, implicit disease inferences, and patient emotional tendencies, and existing speech parsing systems cannot accurately understand these implicit semantics, affecting the accuracy of intelligent diagnosis and treatment. In addition, medical audio data is often associated with multi-modal data such as patients' imaging data (such as X-rays, CT scans) and electronic medical record texts, but existing audio parsing technologies have not fully integrated this information, resulting in diagnostic results lacking complete context information. The parsing results usually use professional medical terms, such as "high blood sugar index is on the high side" or "β-blockers are recommended for use", which are difficult for ordinary patients to understand, affecting the efficiency of doctor-patient communication and the intelligent diagnosis and treatment experience.
[0005] In the financial field, financial audio parsing technology is mainly used in scenarios such as market analysis, investor conference call records, and customer service interactions. However, existing technologies are still limited to basic speech transcription and keyword recognition, and it is difficult to deeply analyze complex semantics such as market trends and risk assessments in financial speech. For example, in an investor conference call, the speech of executives may contain implicit market expectations, but existing parsing technologies cannot comprehensively analyze information such as financial report texts and market data, resulting in limited accuracy in predicting market sentiment. In addition, financial data is usually associated with information such as contract texts and market trend charts, but existing technologies have not effectively integrated these data, resulting in the lack of in-depth understanding of the market background in the speech parsing results. At the same time, the parsing results are often presented in financial jargon, such as "the intensification of liquidity risk" and "the rise of the market sentiment index", which are difficult for non-professional investors to directly understand and apply to investment decisions. Summary of the Invention
[0006] The main purpose of the present invention is to provide an audio parsing method, device, equipment and storage medium based on feature fusion, aiming to solve the technical problems that existing audio parsing technologies have deficiencies in feature fusion, multimodal information integration and knowledge association, resulting in the lack of in-depth semantic understanding and context relevance in parsing results.
[0007] To achieve the above object, the present invention provides an audio parsing method based on feature fusion, including:
[0008] Obtain a target audio in a target field and a target text associated with the target audio;
[0009] Extract the music feature vector of the target audio through an audio analysis module;
[0010] Perform word segmentation processing and semantic analysis on the target text through a text analysis module to generate a text semantic feature vector;
[0011] Fuse the music feature vector and the text semantic feature vector to generate a fusion feature;
[0012] Construct a knowledge graph including knowledge nodes in the target field;
[0013] Input the fusion feature into the knowledge graph for semantic parsing to generate a parsing result.
[0014] Furthermore, to achieve the above object, the present invention provides an audio parsing device based on feature fusion, including:
[0015] A data acquisition module for obtaining a target audio in a target field and a target text associated with the target audio;
[0016] An audio feature extraction module, configured to extract a music feature vector of the target audio through an audio analysis module;
[0017] A text processing module, configured to perform word segmentation processing and semantic analysis on the target text through a text analysis module to generate a text semantic feature vector;
[0018] A feature fusion module, configured to fuse the music feature vector and the text semantic feature vector to generate a fusion feature;
[0019] A knowledge graph construction module, configured to construct a knowledge graph including knowledge nodes in the target domain;
[0020] A semantic parsing module, configured to input the fusion feature into the knowledge graph for semantic parsing to generate a parsing result.
[0021] Further, to achieve the above object, the present invention further provides a computer device, which includes a memory, a processor, and an audio parsing program based on feature fusion stored in the memory and executable on the processor. When the audio parsing program based on feature fusion is executed by the processor, the steps of the audio parsing method based on feature fusion as described above are implemented.
[0022] Further, to achieve the above object, the present invention further provides a computer-readable storage medium, on which an audio parsing program based on feature fusion is stored. When the audio parsing program based on feature fusion is executed by a processor, the steps of the audio parsing method based on feature fusion as described above are implemented.
[0023] Beneficial effects: The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as medical health, fintech, and culture and art. It discloses an audio parsing method based on feature fusion, including: obtaining a target audio and its associated target text in the target domain, extracting a music feature vector of the target audio, extracting a text semantic feature vector of the target text, fusing the music feature vector and the text semantic feature vector to generate a fusion feature, constructing a knowledge graph including knowledge nodes in the target domain, and inputting the fusion feature into the knowledge graph for semantic parsing to generate a parsing result. By fusing audio and text features and combining with a knowledge graph for in-depth semantic parsing, the present invention realizes multimodal understanding of audio data, improves the accuracy and interpretability of the parsing result, and enhances the relevance between the parsing result and industry knowledge, thereby improving the applicability of the audio parsing technology in fields such as culture and art, medical health, and finance. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:
[0025] Figure 1Schematic diagram of an application environment of an audio parsing method based on feature fusion in an embodiment of the present invention;
[0026] Figure 2 Flowchart of an embodiment of an audio parsing method based on feature fusion of the present invention;
[0027] Figure 3 Schematic diagram of functional modules of a preferred embodiment of an audio parsing device based on feature fusion of the present invention;
[0028] Figure 4 Schematic diagram of a structure of a computer device in an embodiment of the present invention;
[0029] Figure 5 Another schematic diagram of a structure of a computer device in an embodiment of the present invention. Detailed implementation manners
[0030] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0031] The audio parsing method based on feature fusion provided by the embodiments of the present invention can be applied in an application environment such as Figure 1 , where the client communicates with the server through a network. The server can obtain the target audio and its associated target text in the target field through the client, extract the music feature vector of the target audio, extract the text semantic feature vector of the target text, fuse the music feature vector and the text semantic feature vector to generate a fusion feature, construct a knowledge graph containing knowledge nodes in the target field, input the fusion feature into the knowledge graph for semantic parsing, and generate a parsing result. The present invention realizes the multimodal understanding of audio data by fusing audio and text features, combining with a knowledge graph for in-depth semantic parsing, improves the accuracy and interpretability of the parsing result, and enhances the relevance between the parsing result and industry knowledge, thereby improving the applicability of audio parsing technology in fields such as culture and art, medical and health, and finance. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.
[0032] Please refer to Figure 2 , Figure 2 which is a flowchart of an embodiment of an audio parsing method based on feature fusion provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.
[0033] Such as Figure 2As shown in the figure, the audio parsing method based on feature fusion proposed by the present invention includes the following steps:
[0034] S10, obtaining a target audio in a target domain and a target text associated with the target audio;
[0035] In this embodiment, to obtain a target audio in a target domain and a target text associated with the target audio, it is necessary to collect, store, and manage data through various methods to ensure the integrity, accuracy, and availability of audio and text data. In practical applications, the collection of target audio can be completed based on existing data storage systems, real-time audio streams, or audio acquisition devices for specific tasks. The acquisition methods of target audio can include retrieving from an audio database, obtaining from an online media platform, or real-time acquisition through devices such as microphones. For stored audio files, format parsing is required to extract metadata information, including audio format, sampling rate, encoding method, duration, etc., to ensure data compatibility in subsequent processing.
[0036] The target text associated with the target audio generally refers to corpus data related to the audio content, which can be the subtitle, transcription text, lyrics, scriptures, dialogue script, etc. of the audio. The methods for obtaining text include directly retrieving from an existing text database, transcribing from the target audio through automatic speech recognition (ASR) technology, extracting subtitle information from an associated video file, or matching relevant text data from an industry knowledge graph or relevant field literature. The quality of text data is crucial for subsequent semantic analysis, so data cleaning and preprocessing are required, including removing noise characters, correcting spelling mistakes, and standardizing text formats. In addition, when obtaining text, the corresponding relationship between multi-modal data needs to be considered to ensure that the text is aligned with the audio content in the time dimension for subsequent feature fusion and semantic parsing.
[0037] During the data acquisition process, a data annotation and metadata management mechanism can be introduced to provide structured information for audio and text data. For example, audio files can be given category labels, sentiment annotations, source information, etc. to assist in subsequent knowledge graph construction and semantic analysis. Similarly, text data can be made more parsable and semantically relevant through entity recognition, keyword extraction, etc. In terms of data storage, a structured database or a distributed storage system can be used to support the efficient retrieval and management of large-scale audio and text data.
[0038] The matching method of data can be based on timestamp matching, text semantic matching, or manual annotation. During the process of audio and text matching, techniques such as Dynamic Time Warping (DTW) and semantic embedding vector matching can be adopted to ensure the precise alignment of multi-modal data. For large-scale data, using machine learning methods to automatically identify and match audio and text data can reduce manual intervention and improve the automation level of the system.
[0039] In different application scenarios, the ways to obtain audio and text can vary. For example, in the field of healthcare, audio can be obtained from sources such as doctors' consultation recordings, patients' medical record reports, and medical interview programs, and electronic medical records, health documents, etc. can be combined as associated text data. In the financial field, audio can be obtained from sources such as financial news broadcasts, investor conference calls, and market analysis dialogues, and relevant market reports, financial analysis articles, legal and regulatory documents, etc. can be combined as text data. In the field of culture and art, audio data can come from music works, speech recordings, drama dialogues, etc., while text data can be lyrics, scripts, literature materials, etc. In the field of Buddhism, the analysis of audio data usually involves various forms of audio content such as Buddhist music, chanting, and Dharma assembly rituals. To obtain the target audio and its associated text, a multi-modal data collection method for Buddhist cultural content needs to be established to ensure the semantic consistency of audio and text data and improve the analysis ability of Buddhist content. The audio data can include temple chanting recordings, Buddhist ritual music, Dharma talk recordings, etc., while the text data can be corresponding scriptures, ritual explanations, Buddhist canonical interpretations, etc.
[0040] In the implementation of data collection technology, a cloud data collection solution can be used to collect and process audio and text data in real time through a cloud computing platform. For example, in the field of healthcare, cloud-based voice collection devices based on 5G networks can upload the conversations between doctors and patients to the cloud database in real time and automatically match the corresponding electronic medical record texts, improving the integration efficiency of medical information. In the financial field, natural language processing technology can be used to transcribe the audio of financial news broadcasts into text and build a dynamic corpus in combination with real-time financial data streams to provide support for financial decision-making.
[0041] An intelligent data filtering mechanism can be adopted to improve the accuracy of data matching. For example, in a medical scenario, a medical knowledge graph can be used to correct the speech transcription text to eliminate errors in speech recognition; in the financial field, based on keyword matching and financial sentiment analysis technologies, irrelevant text content can be filtered to ensure that the obtained text data is highly relevant to the target audio. In addition, an automated data augmentation technology can be introduced. For example, when obtaining audio data, data augmentation algorithms can be applied to generate samples in different noise environments to enhance the robustness of the system.
[0042] Example illustration: In the scenario of Buddhist Dharma assemblies, it is possible to obtain the chanting audio in large-scale Dharma assemblies and perform intelligent analysis in combination with the corresponding ritual texts. For example, in the Water and Land Dharma Assembly, by collecting the chanting audio of different sessions and matching the Dharma assembly ritual texts, it is possible to analyze the rhythm changes and tone patterns of different chanting methods, and combine computer vision technology to analyze the action information in the ceremony, forming a multi-modal Dharma assembly analysis system. In the field of Buddhist education, it is possible to obtain the recordings of the master's sermons and perform semantic analysis on the content of the explanations in combination with Buddhist canonical texts, transforming complex Buddhist doctrines into easy-to-understand interpretation texts to improve the popularization and educational effect of Buddhist knowledge.
[0043] In the scenario of personal cultivation, a mobile application can be used to collect the chanting audio of users and match the corresponding scripture texts. By analyzing parameters such as chanting rhythm and intonation changes, personalized cultivation suggestions can be provided. For example, for users practicing the Heart Sutra, it is possible to analyze characteristics such as speech speed, pauses, and tone changes during their chanting process, and combine meditation knowledge to provide targeted meditation and chanting guidance to improve the cultivation experience and concentration. In addition, in the digital research of Buddhist culture, it is possible to match the chanting audio and text data of different Buddhist sects and construct a Buddhist audio dataset in combination with historical background information to support the analysis, inheritance, and protection of Buddhist music.
[0044] By obtaining the target audio and its associated text data and adopting efficient data matching and processing methods, the integrity and accuracy of the data can be ensured, laying a foundation for subsequent feature extraction, feature fusion, and knowledge graph construction. Through data collection schemes in different fields, it is possible to adapt to various industry requirements and enhance the wide applicability of the data. In addition, through technologies such as data cleaning, automated matching, and intelligent annotation, the quality of audio and text data is improved, providing more accurate input for semantic analysis, thereby enhancing the parsing effect and application value of the entire system.
[0045] S20, extract the music feature vector of the target audio through the audio analysis module;
[0046] In this embodiment, the main function of the audio analysis module is to extract the music feature vector from the target audio for subsequent feature fusion and semantic analysis. The process of extracting the music feature vector includes multiple key technical steps, and each step involves different audio signal processing methods and feature extraction strategies to ensure that the obtained feature information has high representativeness and analyzability.
[0047] First, it is necessary to perform signal preprocessing on the target audio to improve the accuracy and robustness of feature extraction. This process usually includes operations such as denoising, normalization, and sampling rate adjustment. Denoising can use frequency domain filtering methods, such as band-pass filtering, mean filtering, or wavelet transform, to remove environmental noise and non-related interference signals. The normalization operation can standardize the amplitude range of the audio signal to ensure the consistency of the dynamic range between different audio samples, thereby avoiding feature deviation caused by volume differences. Sampling rate adjustment is to ensure that audio data from different sources has the same sampling rate for subsequent feature extraction and calculation.
[0048] Next, it is necessary to perform time-domain framing processing on the audio signal to obtain short-time audio segments suitable for feature extraction. Since the audio signal has time-varying characteristics and cannot be directly processed as a whole, it is necessary to divide the audio into frames of a fixed length and extract features for each frame. Usually, the short-time Fourier transform (STFT) or Mel-frequency cepstral coefficients (MFCC) are used to decompose the audio signal to extract frequency domain and time domain features. When framing, it is necessary to select appropriate frame lengths and frame shifts. For example, common frame lengths are between 20 - 40 ms, and the frame shift is about 10 ms, to ensure that each audio frame can capture sufficient information while ensuring temporal continuity.
[0049] In the feature extraction stage, various deep learning models can be used to analyze the audio signal. For example, a convolutional neural network (CNN) can be used to extract the frequency domain energy distribution features of the audio, identify melody, rhythm, and timbre change patterns. A recurrent neural network (RNN) or long short-term memory network (LSTM) can be used to capture the temporal dependence relationship of the audio signal and extract the evolution features of the melody contour. In addition, the self-attention mechanism can be combined to perform global feature learning on long audio sequences to improve the modeling ability for complex music structures.
[0050] In addition to basic frequency domain and time domain features, higher-level music attribute information, such as timbre, harmony, and rhythm patterns, can also be extracted. Timbre features can be extracted through formant distribution. For example, the Mel-Spectrogram is used to represent the spectral envelope of the audio signal, and principal component analysis (PCA) or autoencoder is combined for dimensionality reduction to reduce redundant information. Harmony features can be extracted through harmonic-inharmonic separation methods to identify the harmonic components and inharmonic noise in the audio. Rhythm patterns can analyze the time domain energy distribution of the audio signal through methods such as wavelet transform and pulse peak detection to extract the rhythm period information of the audio.
[0051] Finally, all the extracted features will be combined and mapped into a unified feature vector space to form a complete music feature vector. Dimensionality reduction methods (such as PCA, t-SNE) can be used to optimize the feature vector to reduce the computational complexity while maintaining the representativeness of the features. This music feature vector will serve as the input for subsequent multi-modal feature fusion and knowledge graph parsing, providing a basis for the in-depth understanding of audio data.
[0052] Example illustration: In the field of healthcare, feature extraction can be performed on speech data in mental health assessments. For example, by analyzing the speech features of patients, such as pitch fluctuations, speech rate, pause frequency, etc., it can assist in the diagnosis of psychological diseases such as depression and autism. Combining audio features and medical record text data, a more comprehensive patient status assessment system can be constructed to improve the intelligent level of mental health management.
[0053] In the financial field, the recorded phone calls of corporate executives can be analyzed to judge the speech emotions and confidence levels of the speakers. For example, in an earnings conference call of a company, the speech features of the CEO or CFO can be extracted, including speech rate, pitch, stress changes, etc., and combined with natural language processing techniques to analyze the speech content to predict market sentiment and investment risks.
[0054] In the field of Buddhism, feature extraction can be performed on the chanting audio of different sects to analyze the changes in pitch, rhythm, and harmony patterns. For example, the chanting audio of Tibetan Buddhism can be analyzed to identify its unique long-note chanting pattern and combined with ritual texts for matching to analyze the functions and characteristics of different ritual music. In addition, the extracted audio features can be associated with the Buddhist cultural knowledge graph to study the evolution of Buddhist music styles in different historical periods and regions, providing technical support for the digital protection of Buddhist culture.
[0055] By extracting the music feature vector of the target audio through the audio analysis module, the core feature information of the audio data can be efficiently obtained, providing high-quality input for subsequent feature fusion and semantic parsing. By combining multiple feature extraction techniques, the parsing ability of the audio content is improved, enabling the system to accurately identify key attributes such as melody, timbre, and rhythm. In addition, by optimizing the feature extraction process, redundant information can be reduced and the computational efficiency can be improved.
[0056] S30, perform word segmentation processing and semantic analysis on the target text through the text analysis module to generate a text semantic feature vector;
[0057] In this embodiment, the main function of the text analysis module is to perform word segmentation and semantic analysis on the target text, so as to extract the semantic information of the text and represent it as a structured feature vector for subsequent multi-modal feature fusion and semantic parsing. This process includes multiple key steps such as text preprocessing, word segmentation, semantic feature extraction, and vectorization representation.
[0058] First, it is necessary to perform text preprocessing on the target text to ensure the quality and consistency of the text data. Text preprocessing usually includes character cleaning, removing meaningless symbols, normalization processing, etc. Character cleaning is to remove invalid characters in the text, such as HTML tags, special symbols, spaces, etc., to ensure the accuracy of subsequent word segmentation and semantic analysis. Normalization processing includes converting traditional Chinese characters to simplified Chinese, unifying different styles of punctuation marks, handling case conversion, etc. In addition, when dealing with multi-language texts (such as Sanskrit, Tibetan, Chinese, etc. in the field of Buddhism), multi-language recognition and language conversion are required to ensure that texts in different languages can be processed using a unified analysis method.
[0059] Next, it is necessary to perform word segmentation on the text to split the continuous text sequence into words or phrases with independent meanings. Word segmentation methods can be divided into rule-based word segmentation, statistical model-based word segmentation, and deep learning-based word segmentation. Rule-based word segmentation methods match through preset dictionaries and rules, such as the maximum matching method (MM), reverse maximum matching method (RMM), etc. Statistical model-based word segmentation methods use statistical learning methods (such as the hidden Markov model HMM) to calculate the probability distribution of words and achieve adaptive word boundary recognition. Deep learning-based word segmentation methods use neural network models, such as BERT, BiLSTM-CRF, etc. The models trained on large-scale corpora can better adapt to text word segmentation in different fields.
[0060] After word segmentation, semantic analysis of the text is required to extract the deep semantic features of the text. The methods of semantic analysis include word vector representation, dependency syntactic analysis, named entity recognition (NER), sentiment analysis, etc. Word vector representation converts the words in the text into high-dimensional vector forms, making words with similar semantics closer in the vector space. Common methods include Word2Vec, GloVe, FastText, etc. In addition, in recent years, the BERT model based on the Transformer architecture can generate context-related word vectors, further improving the accuracy of text semantic representation. Dependency syntactic analysis can identify the grammatical relationships within a sentence, such as the subject-predicate-object structure, modification relationships, etc., to extract more fine-grained semantic information. Named entity recognition (NER) is used to identify specific entities in the text, such as person names, place names, professional terms, etc., which has important application value in specific fields (such as finance, medicine, Buddhism). Sentiment analysis can classify the sentiment tendency of the text content, such as positive, neutral, negative, etc., and is applicable to scenarios such as public opinion monitoring and customer feedback analysis.
[0061] Finally, the extracted text semantic features need to be converted into text semantic feature vectors for subsequent multimodal fusion and knowledge graph construction. The text semantic feature vectors can be represented in different ways, such as TF-IDF vectors, bag-of-words models (BoW), topic models (LDA), deep semantic vectors (BERT Embedding), etc. The text vectorization methods based on deep learning can better capture the context information of the text, enabling the feature vectors to reflect richer semantic information and improving the accuracy of parsing.
[0062] In different application scenarios, the implementation methods of the text analysis module can vary. For example, in the field of healthcare, text data such as electronic medical records, medical papers, and doctor-patient consultation records can be analyzed to extract key information such as patient symptom descriptions, treatment plans, and drug names, and represented as structured text semantic feature vectors. In the financial field, text data such as financial news, corporate financial reports, and market analysis reports can be parsed to extract information such as company names, market indicators, and economic trends to support financial market analysis and risk assessment. In the field of Buddhism, content such as Buddhist scriptures, Dharma assembly rituals, and sermon texts can be parsed to extract key Buddhist concepts, sect information, ritual structures, etc., and combined with audio data for semantic analysis.
[0063] In specific implementation methods, different natural language processing (NLP) technologies can be adopted to optimize the effect of text analysis. For example, in the field of medical and health, medical knowledge graphs can be combined to enhance the semantics of texts. For instance, the Unified Medical Language System (UMLS) can be used to standardize medical terms, improving the accuracy of term matching. In the financial field, a financial entity recognition model (FinBERT) can be combined for in-depth analysis of financial texts to more accurately extract financial information. In the field of Buddhism, a BERT-based text embedding model can be used to understand the context semantics of Buddhist canonical texts and perform semantic reasoning in combination with a Buddhist knowledge base to generate more accurate text semantic feature vectors.
[0064] In addition, in multi-language processing, cross-language models (such as XLM-R, mBERT) can be adopted to support unified semantic analysis of texts in different languages. For example, in Buddhist studies, the Kangyur in Tibetan, the Tripitaka in Chinese, and the Pali Tipitaka can be aligned to extract cross-language Buddhist concepts and perform semantic mapping to support intelligent analysis of multi-language Buddhist content.
[0065] By performing word segmentation and semantic analysis on the target text through the text analysis module, the deep semantic features of the text can be effectively extracted and converted into computable text semantic feature vectors, providing high-quality data support for subsequent multi-modal fusion and knowledge graph parsing. By combining different word segmentation methods, semantic analysis technologies, and text vectorization models, the accuracy and adaptability of text features are improved, enabling the system to process text content in different fields. In addition, by introducing cross-language models and knowledge enhancement technologies, the ability of text parsing can be expanded to make it applicable to multi-language and multi-field application scenarios, improving the generality and scalability of parsing results.
[0066] S40. Fuse the music feature vector and the text semantic feature vector to generate a fused feature;
[0067] In this embodiment, the process of fusing the music feature vector and the text semantic feature vector aims to construct a unified multi-modal feature representation for subsequent knowledge graph matching and semantic parsing. Due to the differences in data structure and feature space between music and text, specific methods are needed to align and fuse the information of the two modalities to obtain a fused feature that can fully express the association between audio and text.
[0068] First, the music feature vector is extracted from the target audio and usually contains feature information such as melody, rhythm, timbre, and harmony. The text semantic feature vector is extracted from the target text and includes information such as word vectors, semantic relations, and syntactic structures. Since the original representation forms of these two features are different, methods such as dimensionality reduction, feature normalization, and alignment are needed to make them comparable in the same feature space. For example, principal component analysis (PCA) or autoencoder can be used to reduce the dimensionality of the high-dimensional music features to make their dimensions consistent with those of the text semantic features.
[0069] In terms of the fusion method, cross-modal attention mechanism (Cross-Modal Attention) or self-supervised learning method is mainly used to enhance the correlation between the music feature vector and the text semantic feature vector. First, calculate the cosine similarity matrix between the two to judge the similarity degree between the music feature and the text feature in the semantic space. Then, use the linear projection method to map the music feature vector into a query vector, map the text semantic feature vector into a key vector and a value vector, and calculate the attention weights of the music feature and the text feature through the multi-head attention mechanism (Multi-Head Attention) to generate cross-modal correlation weights.
[0070] The construction of the fused features can adopt the weighted concatenation (Weighted Concatenation) or adaptive fusion (Adaptive Fusion) method. The weighted concatenation method directly connects the music feature vector and the text semantic feature vector and adjusts the influence of different modal features according to the attention weights. The adaptive fusion method learns the optimal combination method of different modalities through a neural network to dynamically adjust the weight ratio of music and text, so that the fused features can more accurately represent the relationship between the audio content and the text semantics. Finally, the fused features can be further processed for dimensionality reduction to reduce the computational complexity while maintaining their semantic integrity.
[0071] In different application scenarios, the implementation methods of fusing the music feature vector and the text semantic feature vector can be different. For example, in the field of medical and health, the fusion of doctors' inquiry voices and patients' medical record texts needs to emphasize the relationship between the disease description and the voice features such as the doctor's intonation and speaking speed. Therefore, an attention mechanism based on semantic alignment can be adopted to ensure that the text semantics of the disease description and the doctor's voice style work together to improve the accuracy of the diagnosis suggestions.
[0072] In the financial field, the integration of investors' conference call audio and corporate announcement texts requires attention to the matching of changes in the speaker's tone (such as emphasizing certain keywords, pauses in speech, etc.) with the news content. Therefore, an adaptive fusion method based on time series can be adopted to analyze the synchronization of audio emotion features and text market emotion descriptions to judge the degree of influence of market emotion changes.
[0073] In the field of Buddhism, the integration of chanting audio and scripture texts requires attention to the combination of chanting rhythm, intonation and text semantics. Therefore, a method based on melody-text semantic alignment can be adopted to analyze the melody pattern of chanting and the structure of scriptures to match the chanting styles of different rituals. For example, the long-tone chanting method of Tibetan Buddhism matches a specific scripture structure, while the short chanting style of Chan Buddhism may correspond to different semantic structures. By integrating melody and text semantic information, the accuracy of chanting analysis can be improved and applied to scenarios such as Buddhist teaching and cultural dissemination.
[0074] By fusing music feature vectors and text semantic feature vectors, the ability to deeply analyze audio content can be enhanced, enabling the system to not only understand low-level features such as the melody and rhythm of the audio, but also combine text semantics to understand the cultural, industry or emotional information behind it. This fusion method improves the accuracy, interpretability and applicability of audio analysis.
[0075] S50, construct a knowledge graph containing knowledge nodes in the target domain;
[0076] In this embodiment, the process of constructing a knowledge graph containing knowledge nodes in the target domain mainly involves five key steps: data collection, generation of knowledge candidate nodes, semantic annotation, construction of hierarchical structure, and generation of knowledge graph. The core goal of this process is to establish a knowledge base for subsequent semantic parsing and reasoning by structuring multi-modal data (including audio features, text semantic features, etc.) in the target domain.
[0077] First, it is necessary to collect data resources in the target domain, including audio data, text data, and structured data. Audio data usually includes music features of the target audio, such as melody, rhythm, timbre, etc., while text data includes associated text corpora, such as lyrics, chanting texts, medical records, financial news, etc. In addition, structured data may come from industry databases, existing knowledge bases (such as UMLS for healthcare, EDGAR database for financial data, Tripitaka database for Buddhism, etc.), and these data can provide reliable entity information for the construction of the knowledge graph.
[0078] After collecting data, it is necessary to identify and generate a preliminary set of knowledge candidate nodes. Knowledge candidate nodes are usually composed of core concepts, objects, events, etc., which can reflect the basic knowledge system of the target domain. For example, in the medical field, the nodes can be "diseases", "drugs", "treatment plans"; in the financial field, the nodes can be "market trends", "corporate financial reports", "investment strategies"; in the Buddhist field, the nodes can be "scriptures", "cultivation methods", "sects". Candidate nodes can be automatically generated through methods such as keyword extraction, named entity recognition (NER), and dependency syntax analysis.
[0079] Next, it is necessary to perform semantic annotation on the knowledge candidate nodes to ensure that the attributes, definitions, and relationships of each node are clear. Semantic annotation involves assigning information such as names, classifications, definitions, and attributes to the nodes. For example, in the medical field, the attributes that a "cold" node may contain are "symptoms: runny nose, headache, cough", "treatment methods: antiviral drugs, rest", "associated diseases: influenza". In the financial field, the attributes that a "market volatility" node may contain are "influencing factors: policy changes, investor sentiment", "market types: stock market, foreign exchange market", etc. In the Buddhist field, the "meditation" node can contain "cultivation methods: sitting meditation, chanting", "classical source: Platform Sutra of the Sixth Patriarch", "goals: enlightenment, concentration".
[0080] After completing the semantic annotation of the nodes, it is necessary to construct the hierarchical structure and association relationships of the knowledge nodes, that is, to determine the subordinate relationships, causal relationships, similarities, etc. between different knowledge nodes. Methods based on graph neural networks (GNNs) or relation extraction models can be used to hierarchically organize the set of knowledge candidate nodes and construct the connections between the nodes. For example, in the medical field, "respiratory diseases" has a hierarchical relationship with "pneumonia" and "asthma", and "pneumonia" has a treatment relationship with "antibiotics". In the financial field, "market volatility" may have a causal relationship with "economic policy adjustment". In the Buddhist field, "Prajna" and "emptiness" may have a similar relationship, and "meditation" and "entering samadhi" may be in a causal relationship.
[0081] Finally, based on the set of knowledge nodes and their relationships, construct the final knowledge graph. The construction of the knowledge graph can use RDF (Resource Description Framework), Neo4j graph database, or knowledge embedding methods based on deep learning to store all the knowledge nodes and relationships in a structured way to ensure that they can be used for semantic parsing and knowledge reasoning.
[0082] By constructing a knowledge graph containing knowledge nodes in the target domain, the semantic understanding ability of audio parsing can be effectively improved, enabling it to perform in-depth semantic reasoning and decision support by combining the existing knowledge system. It can enhance the relevance of cross-modal data, enabling audio, text, and structured data to form a complete knowledge network, and improving the accuracy, interpretability, and industry applicability of parsing.
[0083] S60, input the fusion feature into the knowledge graph for semantic parsing to generate a parsing result.
[0084] In this embodiment, the semantic parsing of the knowledge graph depends on the matching between the fusion feature and the knowledge nodes in the graph, as well as the reasoning process after matching. The fusion feature includes music features such as melody, rhythm, timbre, and harmony extracted from the target audio, and text semantic features such as morphology, syntax, and semantic structure parsed from the target text. Since the data formats of audio and text features are different, a unified representation method is needed to map them to the same semantic space for effective matching.
[0085] The matching process adopts a method based on vector similarity. After inputting the fusion feature into the knowledge graph, the system first calculates its similarity with the existing knowledge nodes in the graph. The calculation methods include cosine similarity, Euclidean distance, dot product similarity, etc., to ensure that the fusion feature can find the most similar knowledge node in the high-dimensional vector space. The nodes in the knowledge graph contain different concepts, such as diseases, drugs, and treatment methods in the medical field, market trends and investment strategies in the financial field, cultivation methods and classic entries in the Buddhist field, etc.
[0086] After matching relevant knowledge nodes, the knowledge graph performs semantic reasoning based on the existing relationship structure to expand the information scope and provide a richer parsing result. The reasoning process is based on methods such as Graph Embedding, Path Search, and Relational Inference, using the hierarchical structure of the knowledge graph and the association relationships between entities to screen out the most relevant knowledge information. For example, in the medical field, if the "high fever" node is matched, the system can find related diseases such as "flu" and "pneumonia" according to the graph relationship, and further infer possible complications or treatment plans in combination with the patient's medical record.
[0087] During the parsing process, the different modal weights of the fusion feature may affect the final reasoning effect. Therefore, an adaptive weighting strategy is needed to dynamically adjust the weights of different modal features during the reasoning process. For example, in the financial field, when audio analysis detects anxiety in an investor conference call and text analysis extracts market information such as "corporate earnings are lower than expected", the system can adjust the weights of text features and audio emotion features according to historical data to ensure that the reasoning result can reflect the actual situation of the market.
[0088] Finally, the parsing results are stored in a structured manner and provided to subsequent semantic conversion or decision support processes. The structured parsing results include core knowledge nodes, inference chains, possible extended information, etc., in order to provide intelligent data analysis and auxiliary decision-making capabilities in different fields.
[0089] Example illustration: In the field of healthcare, during a doctor's consultation, the doctor records the patient's condition. The system obtains this consultation audio and analyzes it in combination with the patient's electronic medical record text. The audio analysis module extracts features from the doctor's speech data, obtaining speech features such as speech rate, pitch, and pauses, and at the same time parses medical terms in the audio, such as "high fever", "cough", and "difficulty breathing". The text analysis module parses the medical record text, identifies the disease descriptions, diagnosis information, and treatment plans therein, and generates text semantic feature vectors. After inputting the fused features of the audio and text into the medical knowledge graph, the system retrieves medical nodes related to the disease, such as "influenza" and "bronchitis", through semantic matching, and further infers possible complications, recommended treatment plans, and drug options. The parsing results are used to provide intelligent diagnostic suggestions for doctors, assist doctors in judging the condition, and improve the standardization of medical record keeping. In addition, the system can convert the parsing results into health advice understandable to patients through an automatic semantic conversion function, improving patients' awareness and compliance with their own conditions.
[0090] In the financial field, the audio of an investor conference call is obtained and analyzed in combination with relevant financial reports and market analysis texts. The audio analysis module extracts the intonation changes and pause features of the speeches of corporate executives, and combines with the text analysis module to parse the key information in corporate financial reports and market news, such as "profit decline", "intensified market volatility", and "policy adjustment". After fusing the audio and text features, they are input into the financial knowledge graph for parsing. The system matches relevant economic indicators in the knowledge graph, such as "market sentiment index" and "industry profit trend", and conducts semantic reasoning in combination with historical data to predict the possibility of corporate stock price fluctuations. The parsing results can be used in a financial analysis system to provide market risk assessment reports for investors, and intelligently generate investment strategy suggestions based on the speech emotions of corporate executives and the market reactions of text data, helping investors optimize their investment decisions.
[0091] In the field of Buddhism, the chanting audio of temple ceremonies or Buddhist lectures is acquired by the system and matched with the corresponding Buddhist scripture texts for analysis. The audio analysis module extracts musical features such as chanting rhythm, pitch changes, and resonance characteristics. At the same time, the text analysis module analyzes the content of the scriptures and extracts the core Buddhist concepts, such as "dependent origination", "emptiness", and "six paramitas". By fusing audio and text features, the system inputs them into the Buddhist knowledge graph to match relevant scriptures and cultivation methods, such as the relationship between "Prajna Paramita" and the Heart Sutra and the Diamond Sutra, and infers the meaning of this chanting and its role in the ritual. The analysis results can be used for intelligent Buddhist teaching, providing easy-to-understand interpretations to help practitioners deeply understand the content of the scriptures. In addition, the system can convert professional Buddhist terms into modern language expressions through a semantic conversion module, facilitating the public's understanding of Buddhist doctrines and improving the dissemination effect of Buddhist culture.
[0092] By inputting the fused features into the knowledge graph for semantic analysis, the audio analysis results can be made more accurate, interpretable, and extended to a wider range of application scenarios. It enhances the relevance of cross-modal data, enabling audio and text information to be deeply integrated with the domain knowledge system, and improving the accuracy, interpretability, and industry applicability of the analysis.
[0093] The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as healthcare, fintech, and culture and art. It discloses an audio analysis method based on feature fusion, including: acquiring a target audio and its associated target text within a target domain, extracting a music feature vector of the target audio, extracting a text semantic feature vector of the target text, fusing the music feature vector and the text semantic feature vector to generate a fused feature, constructing a knowledge graph containing knowledge nodes of the target domain, inputting the fused feature into the knowledge graph for semantic analysis, and generating an analysis result. By fusing audio and text features and combining with a knowledge graph for in-depth semantic analysis, the present invention realizes multi-modal understanding of audio data, improves the accuracy and interpretability of the analysis results, and enhances the relevance of the analysis results to industry knowledge, thereby improving the applicability of audio analysis technology in fields such as culture and art, healthcare, and finance.
[0094] In one embodiment, the above S20 includes:
[0095] S201, performing frame addition and windowing processing on the waveform of the target audio to generate time-domain signal segments;
[0096] S202, extracting frequency-domain energy distribution features from the time-domain signal segments through a convolutional neural network in the audio analysis module;
[0097] S203, extracting cross-frame melody contour evolution features from the time-domain signal segments through a recurrent neural network in the audio analysis module;
[0098] S204, extract the formant distribution features of the target audio as timbre features through the audio analysis module;
[0099] S205, fuse the frequency-domain energy distribution features, melody contour evolution features, and timbre features to generate the music feature vector.
[0100] In this embodiment, the goal of the audio analysis module is to extract meaningful music feature vectors from the target audio for subsequent multimodal fusion and knowledge graph matching. This process involves performing time-domain and frequency-domain analysis on the audio signal and using deep learning models to extract key features, including frequency-domain energy distribution features, melody contour evolution features, and timbre features.
[0101] First, perform frame addition and windowing on the target audio to convert the continuous audio waveform into a series of short-time signal segments. The audio signal is a continuously changing time-series data. Directly analyzing the complete waveform will result in information loss. Therefore, short-time Fourier transform (STFT) is needed to perform framing. Each frame covers a certain length of audio data, and a Hann window or Hamming window is used for windowing to reduce spectral leakage at the frame boundary. This process can ensure the stable extraction of audio features on a short time scale, enabling subsequent analysis to capture local change patterns in the time-series signal.
[0102] Second, extract the frequency-domain energy distribution features of the framed audio signal through a convolutional neural network (CNN). CNN is highly sensitive to local patterns (such as spectral peaks, formants, etc.), so it is suitable for analyzing the spectrogram of audio. After the audio signal undergoes short-time Fourier transform, it can be converted into a two-dimensional time-frequency graph (Spectrogram), and then input into the CNN to extract frequency-domain features. These features reflect the overall energy distribution, resonance structure, and energy intensity changes in different frequency bands of the audio. For example, the energy peak in the high-frequency band can characterize the brightness of the timbre, while the energy distribution in the low-frequency band can reflect the heaviness of the bass.
[0103] Then, use a recurrent neural network (RNN), especially the long short-term memory network (LSTM) or gated recurrent unit (GRU), to extract the melody contour evolution features. Different from CNN which mainly focuses on local static patterns, RNN is suitable for processing time-series data, so it can learn the changing rules of audio on the time axis. For example, features such as the trend (rising or falling) of the melody, the speed of the rhythm, and the undulation of musical phrases can be used by LSTM to remember the relationship between previous and subsequent frames, and then extract high-level features reflecting the melody structure. These features are crucial for analyzing the emotional features and styles of music.
[0104] In addition, the system also extracts formant distribution features as timbre features. Formants are the main frequency peaks in an audio signal, which determine the basic characteristics of timbre. Methods for extracting formant features include linear predictive coding (LPC), Mel-frequency cepstral coefficients (MFCC), spectral envelope analysis, etc. The distribution of formants can reflect the sound quality characteristics of audio, such as the timbre of human voices and the discrimination of musical instrument timbres. For example, in chanting audio, formants can help distinguish the chanting styles of different sects, such as the long-tone chanting of Tibetan Buddhism and the short recitation of Han Buddhism.
[0105] Finally, the above frequency-domain energy distribution features, melody contour evolution features, and timbre features are fused to construct a complete music feature vector. The fusion methods include feature concatenation, weighted fusion, or Transformer-based cross-modal feature mapping. The fused feature vector not only retains the overall frequency characteristics of the audio but also can express the dynamic changes of the melody and the personalized attributes of the timbre, making it suitable for subsequent text matching and knowledge graph parsing.
[0106] In this embodiment, through the method of multi-level feature extraction, in-depth analysis of the target audio is achieved, ensuring the complete expression of music features and improving the matching degree between audio content and text semantic information. It can more accurately obtain the melody changes, timbre features, and spectral energy distribution of the audio, making its applications in multiple fields such as medical health, finance, and Buddhism have higher parsing accuracy and interpretability.
[0107] In one embodiment, the above S40 includes:
[0108] S401, fuse the music feature vector and the text semantic feature vector to generate a music-text fusion feature;
[0109] S402, obtain the target scene image associated with the target audio, and extract image features from the target scene image;
[0110] S403, fuse the music-text fusion feature and the image features to generate a multi-modal fusion feature.
[0111] In this embodiment, in the process of fusing music feature vectors and text semantic feature vectors, it is necessary to establish cross-modal correlation relationships so that data from different sources can be effectively mapped in the same feature space and used for subsequent semantic parsing and knowledge reasoning. The music feature vectors come from the target audio and include key features such as melody, rhythm, timbre, and harmony. The text semantic feature vectors come from the target text and contain syntactic structures, lexical relationships, and contextual meanings. Since the data forms and expression methods of the two are different, the system adopts cross-modal feature mapping methods, such as cross-modal conversion based on self-attention mechanisms, or through unified embedding space mapping methods, to convert music features and text features into a unified representation form that can be fused.
[0112] After establishing the initial mapping, the system calculates the similarity matrix of the music feature vectors and text semantic feature vectors to measure their correlation in the semantic space. The calculation methods include cosine similarity, Euclidean distance, or cross-modal alignment strategies based on Transformers to ensure that data in the two modalities can establish effective mappings in the high-dimensional vector space. On this basis, the system uses linear projection methods to map the music feature vectors into query vectors respectively, and map the text semantic feature vectors into key vectors and value vectors, and calculates the weight relationship between the two through the multi-head attention mechanism. This process can capture the deep semantic associations between music features and text features. For example, when analyzing Buddhist chanting audio, it can judge whether a certain melody matches a specific concept in the scriptures.
[0113] After obtaining the music-text fusion features, the system further obtains the target scene image and extracts the key visual features therein. The target scene image usually comes from visual data related to the audio, such as medical images, market data visualization charts, images of Dharma assembly sites, etc. To extract effective features from the target scene image, the system uses deep convolutional neural networks or vision transformers for high-dimensional image feature extraction and feature alignment to ensure that the image information can be jointly modeled with the music-text fusion features.
[0114] After completing the feature extraction, the system needs to further fuse the music-text fusion features with the image features to form multi-modal fusion features. Common fusion methods include feature concatenation, cross-modal attention, or multi-modal representation learning based on Transformers. For example, the system can adopt the self-attention mechanism based on Transformers to make the music-text fusion features and the image features correlate with each other and adjust the weights of different modalities according to the correlation strength to generate complete multi-modal fusion features. Finally, the multi-modal fusion features are used for subsequent semantic parsing and knowledge graph matching, enabling the system to simultaneously utilize audio, text, and image information for deeper semantic understanding and reasoning.
[0115] For example, the fusion features of doctors' consultation audio and patients' electronic medical record texts can be used for disease matching. As target scene images, patients' medical images (such as CT and MRI) can further provide visual diagnostic evidence. The system can analyze the symptom text information described by doctors, and combine it with the patients' medical image data to extract the image features of key lesion areas, form multi-modal fusion features, and finally input them into the medical knowledge graph for analysis to infer possible diseases, such as pneumonia, nodules, etc. In addition, it can be used for medical education to provide medical students with case-based multi-modal data analysis to improve their diagnostic reasoning ability.
[0116] For example, the audio of an investors' conference call is analyzed to extract features such as the tone, speech rate, and pauses of executives' speeches, and combined with the text of market analysis reports to generate music-text fusion features. At the same time, the system extracts image features from market data visualization charts (such as K-line charts and fund flow charts), aligns them with the audio-text fusion features, and finally generates multi-modal fusion features. These features can be used for financial knowledge graph reasoning to predict market trends or stock price fluctuations and provide decision-making suggestions for investors. It can also be used for financial supervision to analyze the tone changes in investors' meetings and combine them with market data visualization information to identify potential market manipulation behaviors.
[0117] For example, the chanting audio in a monastery is analyzed, fused with the corresponding Buddhist scripture text, and combined with the on-site video data of the Dharma assembly to extract visual elements such as Buddha statues, Dharma instruments, and the actions of monks. For example, when analyzing a water and land Dharma assembly, the system can analyze the melodic features of the chanting audio, match the corresponding scripture text, and at the same time combine the on-site image data to extract the visual features of the ritual execution to determine whether the Dharma assembly complies with the ritual norms and provide an intelligent Buddhist analysis report. In addition, through multi-modal data fusion, learners can understand the overall relationship between the chanting styles, scripture semantics, and Dharma assembly rituals of different sects, improving the dissemination effect of Buddhist culture.
[0118] In this embodiment, by fusing music feature vectors, text semantic feature vectors, and target scene image features, the system can provide more accurate semantic analysis capabilities in various application scenarios. It enhances the deep correlation ability of multi-modal data, enabling audio, text, and image information to be jointly used for intelligent reasoning, and improving the accuracy and interpretability of audio analysis.
[0119] In one embodiment, the above S401 includes:
[0120] S4011, determining the cosine similarity matrix between the music feature vector and the text semantic feature vector;
[0121] S4012. Perform linear projection processing on the music feature vector to generate a query vector, and perform linear projection processing on the text semantic feature vector to generate a key vector and a value vector;
[0122] S4013. Based on the query vector, key vector, and value vector, generate an initial attention score through a multi-head attention module;
[0123] S4014. Perform weighted summation of the cosine similarity matrix and the initial attention score to generate a cross-modal association weight;
[0124] S4015. Perform weighted concatenation of the music feature vector and the text semantic feature vector according to the cross-modal association weight, and perform dimensionality reduction processing on the concatenated feature vector to generate the music-text fusion feature.
[0125] In this embodiment, the process of fusing the music feature vector and the text semantic feature vector aims to establish cross-modal feature associations, enabling in-depth fusion of audio data and text data within the same feature space. This process involves similarity calculation, feature mapping, attention mechanism, cross-modal weight calculation, and finally feature dimensionality reduction to ensure that the fused features have optimal information expression capabilities and can be used for subsequent knowledge reasoning and semantic parsing.
[0126] First, calculate the cosine similarity matrix between the music feature vector and the text semantic feature vector. Since music features and text features belong to different modalities, directly performing feature fusion may lead to data mismatch or semantic deviation. Therefore, it is necessary to first calculate their similarity degree. Cosine similarity measures the angle between two vectors in a high-dimensional space, with a value range between -1 and 1. The cosine similarity matrix is used to preliminarily judge the correlation between the music feature vector and the text semantic feature vector for subsequent calculation of cross-modal attention weights.
[0127] Then, perform linear projection on the music feature vector to generate a query vector, and at the same time perform linear projection on the text semantic feature vector to generate a key vector and a value vector. The role of linear projection is to map data of different modalities to the same feature space for matching in the attention mechanism. Common projection methods include transformations based on fully connected layers or dimensionality compression based on low-rank matrix factorization. The query vector is used to extract relevant information, the key vector represents the data storage index, and the value vector represents the actual stored content. Through linear projection, music and text features can be represented in a unified comparable vector form.
[0128] Based on the query vector, key vector, and value vector, the initial attention scores are calculated through the multi-head attention module. Multi-head attention is an improved self-attention mechanism that uses multiple attention heads to assign weights to input features from different angles to enhance the model's expressive power. The attention scores calculated by each attention head are used to measure the degree of association between the music feature vector and the text semantic feature vector, and capture different levels of semantic information within multiple subspaces.
[0129] Next, the cosine similarity matrix and the initial attention scores are weighted and summed to obtain the cross-modal association weights. The cosine similarity matrix provides the static similarity between music and text features, while the attention scores provide the context-based dynamic matching information. By weighting and summing the two, the advantages of both information sources can be retained during the fusion process, improving the expressive power of the fused features.
[0130] Finally, according to the cross-modal association weights, the music feature vector and the text semantic feature vector are weighted and concatenated, and the concatenated feature vector is dimensionally reduced to generate the final music-text fusion features. The way of weighted concatenation can be direct concatenation, weighted average, or a combination method based on a transformation network. The purpose of dimensionality reduction is to reduce redundant information and improve the compactness of the features, making the fused features more efficient in subsequent semantic parsing and knowledge reasoning processes.
[0131] In this embodiment, by providing an efficient cross-modal matching mechanism during the multi-modal data fusion process, the music features and text features can be deeply aligned within the same feature space, ensuring the effective transmission and fusion of information. Through cosine similarity, linear projection, and the multi-head attention mechanism, finer-grained feature matching is achieved, improving the expressive power and interpretability of the fused features.
[0132] In one embodiment, the above S50 includes:
[0133] S501, collect data resources within the target domain, identify the core entities in the data resources, and generate a preliminary set of knowledge candidate nodes based on the core entities;
[0134] S502, perform attribute annotation on each node in the set of knowledge candidate nodes to generate a node attribute set containing node names, definitions, and attributes;
[0135] S503, analyze the logical relationships between the knowledge candidate nodes in the set of knowledge candidate nodes to establish a hierarchical structure and association relationships between the knowledge candidate nodes, and generate a set of node relationship edges;
[0136] S504, based on the node attribute set and the set of node relationship edges, generate a knowledge graph containing knowledge nodes in the target domain through a graph structure construction method.
[0137] In this embodiment, the core of constructing a knowledge graph lies in data collection, entity recognition, relationship construction, and graph structure generation. The establishment of a knowledge graph not only requires extracting core concepts from the data resources of the target domain, but also defining the relationships between nodes to ensure that the knowledge graph can be used for subsequent reasoning and parsing.
[0138] First, collect the data resources within the target domain and identify the core entities. The data sources of the target domain can include structured databases, unstructured texts (such as industry reports, medical record texts, Buddhist scripture texts), audio-visual data, etc. The system processes the data resources through data parsing technologies (such as natural language processing, optical character recognition, etc.) and extracts the core entities with practical significance. The core entities are the basic units for constructing the knowledge graph. For example, in the medical field, the core entities include "diseases", "drugs", "treatment methods", etc.; in the financial field, they include "market trends", "investment strategies", "corporate financial reports", etc.; in the Buddhist field, they include "scripture passages", "ritual rules", "Dharma instruments", etc.
[0139] After identifying the core entities, generate a preliminary set of knowledge candidate nodes. The candidate node set is a set of entities that are mentioned multiple times in the data resources and have a certain semantic stability. For example, in medical texts, if "hypertension" appears frequently in multiple cases and medical papers, it can be used as a knowledge candidate node. To improve the accuracy of entity extraction, methods such as named entity recognition (NER), term frequency-inverse document frequency (TF-IDF), and latent Dirichlet allocation (LDA) can be used to automatically screen the core entities.
[0140] Subsequently, perform attribute annotation on each node in the set of knowledge candidate nodes to form a node attribute set. Node attributes include entity names, definitions, characteristic descriptions, etc. For example, in a medical knowledge graph, the attributes of the disease "diabetes" may include "Definition: A chronic metabolic disease caused by abnormal insulin secretion", "High-risk factors: High blood sugar, obesity", "Treatment methods: Drug treatment, diet management", etc. Attribute annotation can be carried out through manual annotation, rule-based automatic extraction, and deep learning models (such as BERT) for automatic expansion.
[0141] Next, analyze the logical relationships between the nodes in the knowledge candidate node set, construct the hierarchical structure and association relationships, and generate the node relationship edge set. The hierarchical structure is used to express the subordinate relationships between concepts. For example, "hypertension" is a subclass of "cardiovascular disease", and "stock market" is a subclass of "financial market". The association relationship is used to establish the semantic connection between entities. For example, the treatment relationship between "diabetes" and "insulin", and the causal relationship between "decline in corporate profits" and "low market sentiment". To automatically construct relationships, methods such as relation extraction (RE) technology, syntactic analysis, and knowledge embedding models (TransE, TransH, TransR) can be used.
[0142] Finally, based on the node attribute set and the node relationship edge set, use the graph structure construction method to generate a complete knowledge graph. The storage and management of the knowledge graph usually adopt graph databases (such as Neo4j) or knowledge graph frameworks (such as RDF / OWL). The generated knowledge graph can support queries, semantic reasoning, and perform intelligent analysis in combination with multimodal data (such as audio, images).
[0143] In this embodiment, by constructing a knowledge graph, the scattered data in the target domain can be structured and associated, and deep intelligent analysis can be provided through semantic reasoning. It enhances the queryability and interpretability of knowledge, supports multimodal data fusion, and improves the accuracy of semantic parsing and knowledge retrieval.
[0144] In one embodiment, the above S60 includes:
[0145] S601, retrieve the target knowledge nodes in the knowledge graph that match the fusion feature, and establish the matching relationship between the fusion feature and the target knowledge nodes to generate a preliminary matching result;
[0146] S602, according to the hierarchical structure and association relationships of the nodes in the knowledge graph, screen the preliminary matching result to determine the node subset corresponding to the fusion feature;
[0147] S603, based on the attribute information of the fusion feature and each node in the node subset, perform semantic reasoning to generate semantic parsing intermediate information;
[0148] S604, integrate the semantic parsing intermediate information, and combine the association relationships between the nodes in the knowledge graph to generate the final parsing result.
[0149] In this embodiment, the process of inputting the fusion feature into the knowledge graph for semantic parsing mainly involves feature matching, hierarchical structure screening, semantic reasoning, and final parsing generation, to ensure that the multimodal fusion data can establish a corresponding relationship with the structured knowledge in the knowledge graph and be used for deep semantic reasoning.
[0150] First, retrieve target knowledge nodes in the knowledge graph that match the fused features, and establish a matching relationship between the two to generate a preliminary matching result. Since the fused features are multi-modal data (including music features, text semantic features, and possibly image features), the matching process needs to adopt cross-modal retrieval methods. For example, vector search, graph embedding matching, or semantic similarity search can be used to ensure a high enough matching degree between the fused features and the nodes in the knowledge graph. For example, in the field of healthcare, if the fused features include the symptoms described by the patient's voice and the electronic medical record text, the matching process can retrieve the disease nodes in the knowledge graph and find the most similar disease categories, such as "flu" or "asthma".
[0151] After completing the preliminary matching, screen the matching results according to the hierarchical structure and association relationships of the knowledge graph to determine the subset of nodes corresponding to the fused features. The knowledge graph usually has hierarchical relationships. For example, "hypertension" belongs to the category of "cardiovascular diseases", and "short-term trading" belongs to the category of "stock investment strategies". Therefore, after matching the preliminary nodes, the hierarchical relationship can be further combined to screen more specific targets. For example, in the field of Buddhism, if the fused features match the node of "meditation", the system can further screen the specific cultivation methods under it, such as "mindfulness meditation" and "silent illumination meditation", to refine the analysis results.
[0152] Next, perform semantic reasoning based on the attribute information of the fused features and the node subset to generate intermediate semantic analysis information. The goal of semantic reasoning is to derive deeper semantic relationships based on the existing structure of the knowledge graph and the information of the fused features. For example, in the financial field, if the fused features indicate that the corporate executives mentioned "weak market demand" in an investor meeting, and the historical data in the knowledge graph shows that such situations usually lead to "profit decline", the system can infer that the company's future profits may be affected. The methods of semantic reasoning can include rule-based reasoning, Bayesian network inference, or knowledge graph completion based on deep learning.
[0153] Finally, integrate the intermediate information of semantic parsing and combine the node association relationships in the knowledge graph to generate the final parsing result. In this process, relation prediction, knowledge graph reasoning algorithms (such as TransE, TransH, ComplEx), or pre-trained knowledge models based on Transformer (such as BERT-KG, GraphBERT) can be used to ensure that the final parsing result can comprehensively reflect the semantics implied by the fusion features. For example, in the medical field, this process can be used to synthesize the patient's symptom descriptions, medical records, and medical literature, infer possible diagnostic conclusions, and provide relevant treatment suggestions.
[0154] For example, in intelligent medical diagnosis, the doctor's inquiring audio and medical record text are fused and then input into the knowledge graph for semantic parsing. The system first retrieves disease nodes in the knowledge graph that match the fusion features, such as "upper respiratory tract infection" or "pneumonia". Then, according to the hierarchical relationship of diseases, it further screens out possible specific disease types, such as "viral pneumonia" or "bacterial pneumonia". Next, combining the patient's specific symptoms (such as coughing, fever), the system performs semantic reasoning to determine which disease the patient may have and generates preliminary diagnostic suggestions, such as "suggest antiviral treatment" or "suggest chest CT examination". Finally, the system integrates all the reasoning information and combines the causes, treatment plans, and drug recommendations in the knowledge graph to generate complete intelligent diagnosis and treatment suggestions.
[0155] In market sentiment analysis, the audio of investors' conference calls, financial report texts, and market data are fused and then input into the knowledge graph for semantic parsing. The system first retrieves financial concepts that match the fusion features, such as "market volatility" and "decline in investment confidence". Then, combining the hierarchical relationship of the financial knowledge graph, it screens out factors more relevant to the current market situation, such as "rising inflation expectations" or "tightening monetary policy". Subsequently, based on the influencing factors of historical market sentiment, the system infers the possible future trends of the market, such as "the stock market will experience more volatility in the short term" or "tech stocks may face valuation adjustments". Finally, the system integrates all the analysis information and combines the economic cycle, macro policies, and industry trends in the knowledge graph to generate investment suggestions, such as "pay attention to safe-haven assets" or "reduce the allocation of high-valuation growth stocks".
[0156] In intelligent Buddhist scripture analysis, after the chanting audio and text analysis features are input into the knowledge graph, the system first retrieves Buddhist knowledge nodes that match the chanting content, such as "Heart Sutra of Prajnaparamita" or "Nirvana in Stillness". Then, based on the hierarchical structure of the knowledge graph, more specific concepts are filtered out, such as "wisdom of emptiness" or "giving rise to a mind without abiding". Combining the melody of the chanting and the text content, the system conducts semantic reasoning, such as "Does the core idea of this passage of scripture conform to the cultivation methods of Chan Buddhism?" or "Does this ritual conform to the requirements of a specific Dharma assembly?". Finally, the system integrates the analysis information, combines the Buddhist doctrines, classical sources, and ritual specifications in the knowledge graph, and generates an intelligent scripture analysis report, such as "This chanting passage is suitable for the opening stage of the Water and Land Dharma Assembly" or "It is recommended to match a certain chapter of the Diamond Sutra for a deeper understanding".
[0157] In this embodiment, through cross-modal matching, hierarchical screening, semantic reasoning, and knowledge association integration, the depth and accuracy of semantic analysis are improved, enabling the audio, text, and knowledge graph to work together to provide accurate, efficient, and intelligent analysis capabilities.
[0158] In one embodiment, after S60 above, it further includes:
[0159] S701, identifying the professional terms included in the analysis result to generate a list of professional term identifiers;
[0160] S702, retrieving the corresponding popular expressions from a preset corpus of popular expressions according to the list of professional term identifiers to generate a mapping of the correspondence between professional terms and popular expressions;
[0161] S703, performing a matching process on the professional terms in the analysis result and the correspondence mapping through a semantic conversion module to generate an intermediate conversion result, where the intermediate conversion result includes the matching information of each professional term and the corresponding popular expression;
[0162] S704, performing a replacement operation on the professional terms in the analysis result through the intermediate conversion result to replace the professional terms in the analysis result with the corresponding popular expressions to generate a converted popular expression result.
[0163] In this embodiment, after completing the semantic analysis of the knowledge graph and generating the analysis result, it is also necessary to perform a popularization conversion on the professional terms in the analysis result to improve the readability and applicability of the analysis result, especially in non-professional user scenarios (such as patient medical consultations, ordinary investor market interpretations, and Buddhist popularization education). This process involves term identification, retrieval of popular expressions, semantic matching, and term replacement to ensure that the final output content can maintain professionalism and be easy to understand.
[0164] First, identify the professional terms in the parsing results to generate a list of professional term identifiers. Professional terms are usually domain-specific words defined in the knowledge graph. For example, in the medical field, terms include "high glycemic index" and "β-blocker"; in the financial field, terms include "quantitative easing" and "leverage ratio"; in the Buddhist field, terms include "nirvana tranquility" and "emptiness wisdom". The system uses methods such as named entity recognition (NER), dictionary matching, and dependency syntactic analysis to detect professional terms in the parsing results and store them in the list of professional term identifiers.
[0165] Next, retrieve the corresponding popularized expressions from a pre-set popularization corpus to generate a mapping of the correspondence between professional terms and popularized expressions. The popularization corpus is a pre-constructed term conversion database that contains the mapping between professional terms and their everyday language. For example, in the medical and health field, the popularized expression corresponding to "β-blocker" may be "a drug that helps lower heart rate"; in the financial field, "leverage ratio" can correspond to "the ratio of the amount borrowed to the owner's equity"; in the Buddhist field, "emptiness wisdom" can correspond to "the enlightenment that transcends all attachments". Matching methods can use techniques such as rule matching (based on a term dictionary), BERT-based semantic matching, and knowledge graph embedding (KG Embedding).
[0166] Then, through the semantic conversion module, match the professional terms in the parsing results with the correspondence mapping to generate an intermediate conversion result. The intermediate conversion result records the matching information of each professional term and its popularized expression to ensure the traceability and accuracy of the replacement process. For example, if the parsing result contains "β-blockers can reduce sympathetic excitability", the intermediate conversion result will provide its corresponding replacement option, such as "a drug that helps lower heart rate can reduce the effects of tension and stress".
[0167] Finally, through the intermediate conversion result, perform a replacement operation on the professional terms in the parsing results to generate the final converted popularized expression result. The replacement methods can use direct replacement, context-aware substitution, and neural network-based text generation to ensure that the converted result is both faithful to the original meaning and conforms to the habits of everyday language expression. For example, in the Buddhist field, if the parsing result contains "nirvana tranquility", the system can replace it with "achieve inner peace and liberation".
[0168] For example, in an intelligent medical condition analysis system, after the doctor's inquiry audio and medical record text are parsed by the knowledge graph, they may contain a large number of medical terms. For example, "This patient may need to receive β-blocker treatment to reduce the excitability of the sympathetic nervous system." The system first identifies the professional terms "β-blocker" and "sympathetic nervous system", and retrieves their corresponding expressions in the popularization corpus, such as "drugs that help reduce heart rate" and "the system that controls the body's automatic responses". The system matches the professional terms with the popularized expressions and replaces them in the final output to make the analysis result more understandable, such as "This patient may need to take drugs that help reduce heart rate to reduce the impact of tension and stress."
[0169] In an intelligent investment report system, the original analysis result that an investor may receive contains "Due to the implementation of quantitative easing policies by the Federal Reserve, the market liquidity has increased significantly, leading to an increase in inflation risk." For non-professional investors, the term "quantitative easing" may be difficult to understand. Therefore, the system first identifies this term and finds its corresponding expression in the popularization corpus, such as "The central bank stimulates economic growth by printing money and buying bonds." Finally, the analysis result is converted to "Since the central bank stimulates the economy by printing money and buying bonds, there is more money in the market, which may lead to rising prices."
[0170] In an intelligent Buddhist analysis system, the system may parse out "Practitioners should view all dharmas with a mind of non-attachment and ultimately reach nirvana and stillness." For general learners, "mind of non-attachment" and "nirvana and stillness" may be difficult to understand. Therefore, the system identifies these terms and retrieves their popularized expressions, such as "Let go of attachments and do not pursue gains and losses" and "Achieve inner peace and liberation." Finally, the analysis result is converted to "Practitioners should let go of attachments, not pursue gains and losses, and ultimately achieve inner peace and liberation."
[0171] Through the above steps, this embodiment can automatically convert complex professional terms into easy-to-understand everyday language, improving the applicability of the analysis result without losing the original information. Compared with traditional term explanation methods, it can combine the context and knowledge graph to achieve accurate, traceable, and context-compliant term conversion, thereby optimizing the user experience.
[0172] In one embodiment, an audio analysis device based on feature fusion is provided. The audio analysis device based on feature fusion corresponds one-to-one with the audio analysis method based on feature fusion in the above embodiment. Refer to Figure 3 , Figure 3 is a schematic diagram of the functional modules of a preferred embodiment of the audio analysis device based on feature fusion of the present invention. The data acquisition module 10, the audio feature extraction module 20, the text processing module 30, the feature fusion module 40, the knowledge graph construction module 50, and the semantic analysis module 60. The detailed description of each functional module is as follows:
[0173] The data acquisition module 10 is used to obtain the target audio within the target field and the target text associated with the target audio;
[0174] The audio feature extraction module 20 is used to extract the music feature vector of the target audio through the audio analysis module;
[0175] The text processing module 30 is used to perform word segmentation processing and semantic analysis on the target text through the text analysis module to generate a text semantic feature vector;
[0176] The feature fusion module 40 is used to fuse the music feature vector and the text semantic feature vector to generate a fusion feature;
[0177] The knowledge graph construction module 50 is used to construct a knowledge graph containing knowledge nodes in the target field;
[0178] The semantic parsing module 60 is used to input the fusion feature into the knowledge graph for semantic parsing to generate a parsing result.
[0179] In one embodiment, the audio feature extraction module 20 is specifically used for:
[0180] Perform frame addition and windowing processing on the waveform of the target audio to generate a time-domain signal segment;
[0181] Extract the frequency-domain energy distribution feature from the time-domain signal segment through the convolutional neural network in the audio analysis module;
[0182] Extract the cross-frame melody contour evolution feature from the time-domain signal segment through the recurrent neural network in the audio analysis module;
[0183] Extract the formant distribution feature of the target audio through the audio analysis module as the timbre feature;
[0184] Fuse the frequency-domain energy distribution feature, the melody contour evolution feature and the timbre feature to generate the music feature vector.
[0185] In one embodiment, the feature fusion module 40 is specifically used for:
[0186] Fuse the music feature vector and the text semantic feature vector to generate a music-text fusion feature;
[0187] Obtain the target scene image associated with the target audio, and extract the image feature from the target scene image;
[0188] Fuse the music-text fusion feature and the image feature to generate a multi-modal fusion feature.
[0189] In one embodiment, the feature fusion module 40 is specifically configured to:
[0190] Determine the cosine similarity matrix between the music feature vector and the text semantic feature vector;
[0191] Perform linear projection processing on the music feature vector to generate a query vector, and perform linear projection processing on the text semantic feature vector to generate a key vector and a value vector;
[0192] Based on the query vector, key vector, and value vector, generate an initial attention score through a multi-head attention module;
[0193] Perform weighted summation of the cosine similarity matrix and the initial attention score to generate a cross-modal association weight;
[0194] Perform weighted splicing on the music feature vector and the text semantic feature vector according to the cross-modal association weight, and perform dimensionality reduction processing on the spliced feature vector to generate the music-text fusion feature.
[0195] In one embodiment, the knowledge graph construction module 50 is specifically configured to:
[0196] Collect data resources in the target domain, identify the core entities in the data resources, and generate a preliminary knowledge candidate node set based on the core entities;
[0197] Perform attribute annotation on each node in the knowledge candidate node set to generate a node attribute set including node names, definitions, and attributes;
[0198] Analyze the logical relationships between the knowledge candidate nodes in the knowledge candidate node set to establish a hierarchical structure and association relationships between the knowledge candidate nodes, and generate a node relationship edge set;
[0199] Based on the node attribute set and the node relationship edge set, generate a knowledge graph including knowledge nodes in the target domain through a graph structure construction method.
[0200] In one embodiment, the semantic parsing module 60 is specifically configured to:
[0201] Retrieve target knowledge nodes in the knowledge graph that match the fusion feature, and establish a matching relationship between the fusion feature and the target knowledge nodes to generate a preliminary matching result;
[0202] According to the hierarchical structure and association relationships of the nodes in the knowledge graph, screen the preliminary matching result to determine a node subset corresponding to the fusion feature;
[0203] Perform semantic reasoning based on the fusion feature and the attribute information of each node in the node subset to generate intermediate semantic parsing information;
[0204] Integrate the intermediate semantic parsing information and combine the association relationships between nodes in the knowledge graph to generate a final parsing result.
[0205] In one embodiment, the semantic parsing module 60 is specifically configured to:
[0206] Identify the professional terms included in the parsing result to generate a list of professional term identifiers;
[0207] Retrieve the corresponding popularized expressions from a preset popularized expression corpus according to the list of professional term identifiers to generate a correspondence mapping between professional terms and popularized expressions;
[0208] Perform a matching process on the professional terms in the parsing result and the correspondence mapping through a semantic conversion module to generate an intermediate conversion result, where the intermediate conversion result includes the matching information of each professional term and the corresponding popularized expression;
[0209] Perform a replacement operation on the professional terms in the parsing result through the intermediate conversion result to replace the professional terms in the parsing result with the corresponding popularized expressions to generate a converted popular expression result.
[0210] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 4 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps of the server side of an audio parsing method based on feature fusion.
[0211] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as Figure 5As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the user side of an audio parsing method based on feature fusion
[0212] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are realized:
[0213] Obtain a target audio in a target field and target text associated with the target audio;
[0214] Extract the music feature vector of the target audio through an audio analysis module;
[0215] Perform word segmentation processing and semantic analysis on the target text through a text analysis module to generate a text semantic feature vector;
[0216] Fuse the music feature vector and the text semantic feature vector to generate a fusion feature;
[0217] Construct a knowledge graph containing knowledge nodes in the target field;
[0218] Input the fusion feature into the knowledge graph for semantic parsing to generate a parsing result.
[0219] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the following steps are realized:
[0220] Obtain a target audio in a target field and target text associated with the target audio;
[0221] Extract the music feature vector of the target audio through an audio analysis module;
[0222] Perform word segmentation processing and semantic analysis on the target text through a text analysis module to generate a text semantic feature vector;
[0223] Fuse the music feature vector and the text semantic feature vector to generate a fusion feature;
[0224] Construct a knowledge graph containing knowledge nodes in the target field;
[0225] Input the fusion feature into the knowledge graph for semantic parsing to generate a parsing result.
[0226] It should be noted that for the functions or steps that can be realized by the above computer-readable storage medium or computer device, reference can be made to the relevant descriptions on the server side and the user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0227] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0228] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.
[0229] It should be noted that if there are software tools or components of other companies in the embodiments of this application, they are only used for illustrative introduction and do not represent actual use. The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. An audio analysis method based on feature fusion, characterized in that: The following steps are involved: Acquire target audio in a target domain and target text associated with the target audio; Extracting a music feature vector of the target audio through an audio analysis module; Perform word segmentation and semantic analysis on the target text through a text analysis module to generate a text semantic feature vector; Fusion of the music feature vector and the text semantic feature vector to generate a fusion feature; Construct a knowledge graph containing knowledge nodes in the target domain; The fused features are input into the knowledge graph for semantic analysis to generate analysis results.
2. The audio analysis method based on feature fusion according to claim 1, characterized in that: Extracting the music feature vector of the target audio through an audio analysis module includes: Performing frame-by-frame windowing processing on the waveform of the target audio to generate time-domain signal segments; Extracting frequency domain energy distribution features from the time domain signal segment through a convolutional neural network in the audio analysis module; Extracting cross-frame melody contour evolution features from the time domain signal segment through a recurrent neural network in the audio analysis module; Extracting the formant distribution characteristics of the target audio as timbre characteristics through the audio analysis module; The frequency domain energy distribution feature, melody contour evolution feature and timbre feature are integrated to generate the music feature vector.
3. The audio analysis method based on feature fusion according to claim 1, characterized in that: Fusion of the music feature vector and the text semantic feature vector to generate a fusion feature includes: Fusion of the music feature vector and the text semantic feature vector to generate a music text fusion feature; Acquire a target scene image associated with the target audio, and extract image features from the target scene image; The music text fusion feature and the image feature are fused to generate a multimodal fusion feature.
4. The audio analysis method based on feature fusion according to claim 3, characterized in that: Fusion of the music feature vector and the text semantic feature vector to generate a music text fusion feature includes: Determining a cosine similarity matrix between the music feature vector and the text semantic feature vector; Performing linear projection processing on the music feature vector to generate a query vector, and performing linear projection processing on the text semantic feature vector to generate a key vector and a value vector; Based on the query vector, the key vector and the value vector, generating an initial attention score through a multi-head attention module; Performing a weighted summation of the cosine similarity matrix and the initial attention score to generate a cross-modal association weight; The music feature vector and the text semantic feature vector are weightedly spliced according to the cross-modal association weight, and the spliced feature vector is subjected to dimensionality reduction processing to generate the music text fusion feature.
5. The audio analysis method based on feature fusion according to claim 1, characterized in that: Construct a knowledge graph containing knowledge nodes in the target domain, including: Collect data resources in the target domain, identify core entities in the data resources, and generate a preliminary set of candidate knowledge nodes based on the core entities; Performing attribute labeling on each node in the knowledge candidate node set to generate a node attribute set including node name, definition and attributes; Analyzing the logical relationships between the knowledge candidate nodes in the knowledge candidate node set to establish a hierarchical structure and association relationships between the knowledge candidate nodes and generate a node relationship edge set; Based on the node attribute set and the node relationship edge set, a knowledge graph containing target domain knowledge nodes is generated through a graph structure construction method.
6. The audio analysis method based on feature fusion according to claim 1, characterized in that: The fusion feature is input into the knowledge graph for semantic analysis to generate analysis results, including: Retrieving a target knowledge node matching the fusion feature in the knowledge graph, and establishing a matching relationship between the fusion feature and the target knowledge node to generate a preliminary matching result; According to the hierarchical structure and association relationship of the nodes in the knowledge graph, the preliminary matching results are screened to determine a node subset corresponding to the fusion feature; Perform semantic reasoning based on the fusion feature and the attribute information of each node in the node subset to generate semantic parsing intermediate information; The intermediate information of the semantic analysis is integrated, and the association relationship between the nodes in the knowledge graph is combined to generate the final analysis result.
7. The audio analysis method based on feature fusion according to claim 1, characterized in that: After inputting the fusion features into the knowledge graph for semantic analysis and generating the analysis results, the method further includes: Identify the professional terms contained in the analysis results and generate a professional term identification list; According to the professional term identification list, corresponding popular expressions are retrieved from a preset popular expression corpus to generate a corresponding relationship mapping between professional terms and popular expressions; Matching the professional terms in the parsing result with the corresponding relationship mapping through a semantic conversion module to generate a conversion intermediate result, wherein the conversion intermediate result includes matching information between each professional term and the corresponding popular expression; The professional terms in the parsing result are replaced by the conversion intermediate result, so as to replace the professional terms in the parsing result with corresponding popular expressions, thereby generating a converted popular expression result.
8. An audio analysis device based on feature fusion, characterized in that: The audio analysis device based on feature fusion includes: A data acquisition module, used for acquiring target audio in a target field and target text associated with the target audio; An audio feature extraction module, used for extracting a music feature vector of the target audio through an audio analysis module; A text processing module, used for performing word segmentation and semantic analysis on the target text through a text analysis module to generate a text semantic feature vector; A feature fusion module, used for fusing the music feature vector and the text semantic feature vector to generate a fusion feature; A knowledge graph construction module is used to construct a knowledge graph containing knowledge nodes in the target domain; The semantic parsing module is used to input the fused features into the knowledge graph for semantic parsing and generate parsing results.
9. A computer device, characterized in that: The computer device includes a memory, a processor, and a feature fusion-based audio parsing program stored in the memory and executable on the processor. When the feature fusion-based audio parsing program is executed by the processor, the steps of the feature fusion-based audio parsing method as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The storage medium stores an audio analysis program based on feature fusion, and when the audio analysis program based on feature fusion is executed by the processor, the steps of the audio analysis method based on feature fusion as described in any one of claims 1-7 are implemented.
Citation Information
Cited By
Voice call content intelligent analysis and seat reminding system
CN120581009A
Complementation method for underground water level and building risk knowledge graph
CN120764644A
Judicial scene-oriented multi-modal data fusion method and system
CN120892985A
Intelligent decision-making method and device based on multiple modes, equipment and medium
CN121122266A
Multimodal intelligent decision-making methods, devices, equipment, and media
CN121122266B