Music retrieval methods, music retrieval devices, electronic devices and storage media

By performing word recognition on the target description text and spectral transformation on candidate music, genre representation vectors are obtained. Combined with genre description words and tag data for filtering, the problem of low accuracy in music retrieval in existing technologies is solved, and more efficient music retrieval is achieved.

CN116595216BActive Publication Date: 2026-05-26PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2023-05-19
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing music retrieval methods rely on manually labeled music genre tags, which cannot fully label all music tracks, resulting in low retrieval accuracy.

Method used

By acquiring the target description text, performing word recognition and spectral transformation, and obtaining the genre representation vector of candidate music, the target music is obtained by combining genre description words and tag data for filtering.

Benefits of technology

It improves the accuracy and precision of music retrieval, enabling better identification and filtering of music that meets the needs of the target audience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116595216B_ABST
    Figure CN116595216B_ABST
Patent Text Reader

Abstract

This application provides a music retrieval method, a music retrieval device, an electronic device, and a storage medium, belonging to the field of artificial intelligence technology. The method includes: acquiring target descriptive text and candidate music, wherein the target descriptive text includes the target object's description of the music; performing word recognition on the target descriptive text to obtain genre description words; performing spectral transformation on the candidate music to obtain candidate music spectrum sequences; based on the candidate music spectrum sequences, obtaining candidate music genre representation vectors corresponding to the candidate music; performing genre identification on the candidate music based on the candidate music genre representation vectors to obtain genre tag data for the candidate music; filtering the candidate music based on the genre description words and genre tag data to obtain target music; and feeding the target music back to the target object. This application embodiment can improve the accuracy of music retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a music retrieval method, a music retrieval device, an electronic device, and a storage medium. Background Technology

[0002] In the music retrieval process, most current retrieval methods rely on manually labeled music genre tags to match the descriptions entered by users. This method often cannot fully label all music tracks and has significant limitations. For example, manual labeling is usually done with single tags, which cannot reflect all genre characteristics of music, leading to low accuracy in music retrieval. Therefore, how to improve the accuracy of music retrieval has become an urgent technical problem to be solved. Summary of the Invention

[0003] The main objective of this application is to provide a music retrieval method, a music retrieval device, an electronic device, and a storage medium, with the aim of improving the accuracy of music retrieval.

[0004] To achieve the above objectives, a first aspect of this application proposes a music retrieval method, the method comprising:

[0005] Obtain target description text and candidate music, wherein the target description text includes the target object's description of the music;

[0006] Word recognition is performed on the target descriptive text to obtain genre description words;

[0007] Perform spectral transformation on the candidate music to obtain a candidate music spectrum sequence;

[0008] Based on the candidate music spectrum sequence, obtain the candidate music genre representation vector corresponding to the candidate music;

[0009] Based on the candidate music genre representation vector, the candidate music is identified into a genre to obtain the genre tag data of the candidate music.

[0010] The candidate music is filtered based on the genre description words and genre tag data to obtain the target music;

[0011] The target music is fed back to the target object.

[0012] In some embodiments, the step of performing word recognition on the target descriptive text to obtain genre description words includes:

[0013] The target description text is segmented to obtain multiple target description words;

[0014] Entity features are extracted from the target descriptive words to obtain entity word features;

[0015] Based on a preset dictionary and the features of the entity words, the descriptive words of the school of thought are selected from the target descriptive words.

[0016] In some embodiments, the step of filtering the candidate music based on the genre description words and the genre tag data to obtain the target music includes:

[0017] A relevance score is obtained by performing a relevance assessment on the descriptive terms of the genre and the tag data of the genre;

[0018] Based on the relevance score, target tags are selected from the genre tag data, wherein the target tags and the genre description words represent the same music genre;

[0019] Obtain the total number of tags for the target tag;

[0020] The candidate music is filtered based on the total number of tags to obtain the target music.

[0021] In some embodiments, feeding back the target music to the target object includes:

[0022] The target music is sorted based on the total number of tags in the target music to obtain a music search list;

[0023] The music search list is then fed back to the target object.

[0024] In some embodiments, performing spectral transformation on the candidate music to obtain a candidate music spectral sequence includes:

[0025] The candidate music is subjected to spectral transformation based on a preset function to obtain the Mel cepstral features of the candidate music;

[0026] The Mel cepstral features are segmented to obtain multiple candidate spectral segments;

[0027] The candidate spectral segment is dimensionality reduced to obtain the target spectral segment;

[0028] The target spectrum segment is spliced ​​together to obtain the candidate music spectrum sequence.

[0029] In some embodiments, obtaining the candidate music genre representation vector corresponding to the candidate music based on the candidate music spectrum sequence includes:

[0030] Based on a preset coding network, feature extraction is performed on the candidate music spectrum sequence to obtain a preliminary music genre representation vector, wherein the coding network includes at least two transformer encoders;

[0031] The preliminary music genre representation vector is fine-tuned to obtain the candidate music genre representation vector.

[0032] In some embodiments, the step of identifying the genre of the candidate music based on the candidate music genre representation vector to obtain the genre label data of the candidate music includes:

[0033] The candidate music is scored based on the preset candidate genre categories and the candidate music genre representation vectors to obtain a music genre score.

[0034] The candidate genre categories are filtered based on the music genre scores to obtain multiple target genre tags for the candidate music.

[0035] Based on the target genre tag, the genre tag data is obtained.

[0036] To achieve the above objectives, a second aspect of this application provides a music retrieval device, the device comprising:

[0037] The acquisition module is used to acquire target description text and candidate music, wherein the target description text includes the description content of the music by the target object;

[0038] The word recognition module is used to perform word recognition on the target description text to obtain genre description words;

[0039] The spectrum transformation module is used to perform spectrum transformation on the candidate music to obtain a candidate music spectrum sequence;

[0040] The genre representation extraction module is used to obtain the candidate music genre representation vector corresponding to the candidate music based on the candidate music spectrum sequence.

[0041] The genre identification module is used to identify the genre of the candidate music based on the candidate music genre representation vector, and obtain the genre tag data of the candidate music.

[0042] The music filtering module is used to filter the candidate music based on the genre description words and the genre tag data to obtain the target music;

[0043] The result feedback module is used to feed back the target music to the target object.

[0044] To achieve the above objectives, a third aspect of the present application provides an electronic device, the electronic device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method described in the first aspect.

[0045] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.

[0046] The music retrieval method, device, electronic equipment, and storage medium proposed in this application acquire target descriptive text and candidate music. The target descriptive text includes the target object's description of the music. Word recognition is performed on the target descriptive text to obtain genre description terms, which facilitates the identification of vocabulary describing music genres within the target descriptive text. This allows for the retrieval of music that meets the requirements from the candidate music based on the genre description terms. Furthermore, spectral transformation is performed on the candidate music to obtain a candidate music spectrum sequence, which facilitates the extraction of audio feature information from the candidate music. Further, based on the candidate music spectrum sequence, a candidate music genre representation vector is obtained, enabling fine-grained representation of the genre information of the candidate music and improving the accuracy of music retrieval. Furthermore, the candidate music genre is identified based on the candidate music genre representation vector to obtain the genre tag data of the candidate music. The candidate music is then filtered based on the genre description words and genre tag data to obtain the target music. The target music is then fed back to the target object. This method can conveniently filter the target music from multiple candidate music based on genre description words to select the target music whose genre category better matches the target object's search needs. It can effectively reduce the amount of search data and improve the accuracy of music search. Attached Figure Description

[0047] Figure 1 This is a flowchart of the music retrieval method provided in the embodiments of this application;

[0048] Figure 2 yes Figure 1 The flowchart of step S102 in the document;

[0049] Figure 3 yes Figure 1 The flowchart of step S103 in the process;

[0050] Figure 4 yes Figure 1 The flowchart of step S104 in the process;

[0051] Figure 5 yes Figure 1The flowchart of step S105 in the process;

[0052] Figure 6 yes Figure 1 The flowchart of step S106 in the process;

[0053] Figure 7 yes Figure 1 The flowchart of step S107 in the process;

[0054] Figure 8 This is a schematic diagram of the structure of the music retrieval device provided in the embodiments of this application;

[0055] Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0056] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0057] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0059] First, let's analyze some of the terms used in this application:

[0060] Artificial intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.

[0061] Natural Language Processing (NLP): NLP uses computers to process, understand, and utilize human language (such as Chinese and English). NLP is a branch of artificial intelligence and an interdisciplinary field of computer science and linguistics, often referred to as computational linguistics. NLP includes syntactic analysis, semantic analysis, and discourse understanding. It is commonly used in machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, intent recognition, information extraction and filtering, text classification and clustering, sentiment analysis, and opinion mining. It involves data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research, and linguistic research related to language computation.

[0062] Information Extraction (NER) is a text processing technique that extracts factual information such as entities, relationships, and events from natural language text and outputs it as structured data. Information extraction is a technique for extracting specific information from text data. Text data is composed of specific units, such as sentences, paragraphs, and chapters. Text information is composed of smaller, specific units, such as characters, words, phrases, sentences, paragraphs, or combinations of these units. Extracting noun phrases, names of people, and place names from text data is an example of text information extraction. Of course, text information extraction techniques can extract information of various types.

[0063] Mel-Frequency Cipstal Coefficients (MFCCs) are a set of key coefficients used to construct a Mel-Frequency Cipstal spectrum. From a segment of a music signal, a set of cepstrum values ​​can be obtained that is sufficient to represent the music signal. The Mel-Frequency Cipstal Coefficients are the cepstrum values ​​derived from this cepstrum (i.e., the spectrum of the spectrum). Unlike a regular cepstrum, the most distinctive feature of the Mel-Frequency Cipstrum is that its frequency bands are uniformly distributed across the Mel scale. In other words, compared to the linear cepstrum representations commonly seen, this frequency band is closer to the non-linear human auditory system. For example, Mel-Frequency Cipstals are frequently used in audio compression techniques.

[0064] Hidden Markov Models (HMMs) are statistical models used to describe a Markov process with hidden, unknown parameters. The challenge lies in determining these hidden parameters from the observable parameters. These parameters are then used for further analysis, such as pattern recognition. In simple Markov models (like Markov chains), the states are directly observable, so the state transition probabilities are the only parameters. In HMMs, the states are not directly observable, but the outputs, depending on those states, are observable. Each state has a possible probability distribution through possible output tokens. Therefore, generating a sequence of tokens through an HMM provides some sequential information about the states. Note that "hidden" refers to the sequence of states passed through the model, not the model's parameters; even if these parameters were precisely known, we still call the model a "hidden" Markov model. HMMs are known for their temporal pattern recognition applications, such as speech, handwriting, gesture recognition, word class tagging, musical notation, partial discharge analysis, and bioinformatics applications.

[0065] Softmax function: The Softmax function is a normalized exponential function.

[0066] Multi-head attention utilizes multiple queries to compute multiple pieces of information from the input in parallel. Each attention focuses on a different part of the input. Hard attention, on the other hand, is based on the expectation of all input information according to the attention distribution.

[0067] Encoding: This involves transforming an input sequence into a vector of fixed length.

[0068] Decoding: This involves transforming a previously generated fixed vector into an output sequence; the input sequence can be text, speech, image, or video; the output sequence can be text or image.

[0069] In the music retrieval process, most current retrieval methods rely on manually labeled music genre tags to match the descriptions entered by users. This method often cannot fully label all music tracks and has significant limitations. For example, manual labeling is usually done with single tags, which cannot reflect all genre characteristics of music, leading to low accuracy in music retrieval. Therefore, how to improve the accuracy of music retrieval has become an urgent technical problem to be solved.

[0070] Based on this, embodiments of this application provide a music retrieval method, a music retrieval device, an electronic device, and a storage medium, aiming to improve the accuracy of music retrieval.

[0071] The music retrieval method, music retrieval device, electronic device, and storage medium provided in this application are specifically described through the following embodiments. First, the music retrieval method in this application embodiment is described.

[0072] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0073] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0074] The music retrieval method provided in this application relates to the field of artificial intelligence technology. The music retrieval method provided in this application can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the music retrieval method, but is not limited to the above forms.

[0075] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0076] Figure 1 This is an optional flowchart of the music retrieval method provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps S101 to S107.

[0077] Step S101: Obtain the target description text and candidate music, wherein the target description text includes the description content of the music by the target object;

[0078] Step S102: Perform word recognition on the target description text to obtain the genre description words;

[0079] Step S103: Perform spectral transformation on the candidate music to obtain the candidate music spectral sequence;

[0080] Step S104: Based on the candidate music spectrum sequence, obtain the candidate music genre representation vector corresponding to the candidate music.

[0081] Step S105: Based on the candidate music genre representation vector, perform genre identification on the candidate music to obtain the genre label data of the candidate music.

[0082] Step S106: Filter candidate music based on genre description words and genre tag data to obtain target music;

[0083] Step S107: Feedback the target music to the target object.

[0084] Steps S101 to S107 of this embodiment involve acquiring target description text and candidate music. The target description text includes the target object's description of the music. Word recognition is performed on the target description text to obtain genre description words, which facilitates the identification of vocabulary describing music genres within the target description text. This allows for the retrieval of suitable music from the candidate music based on the genre description words. Furthermore, spectral transformation is performed on the candidate music to obtain a candidate music spectrum sequence, which facilitates the extraction of audio feature information from the candidate music. Further, based on the candidate music spectrum sequence, a candidate music genre representation vector is obtained, enabling fine-grained representation of the genre information of the candidate music and improving the accuracy of music retrieval. Furthermore, the candidate music genre is identified based on the candidate music genre representation vector to obtain the genre tag data of the candidate music. The candidate music is then filtered based on the genre description words and genre tag data to obtain the target music. The target music is then fed back to the target object. This method can conveniently filter the target music from multiple candidate music based on genre description words to select the target music whose genre category better matches the target object's search needs. It can effectively reduce the amount of search data and improve the accuracy of music search.

[0085] In step S101 of some embodiments, candidate music can be extracted from a preset music database or downloaded from online platforms or other channels. The candidate music can include music of different languages, scenes, and styles. For example, candidate music includes Chinese songs, English songs, or Japanese songs, etc. The target user can input the description of the music they want to obtain through various methods such as typing or voice input. The music retrieval system can identify the description based on a preset content recognition algorithm to obtain the target description text. The target user can be an online user or other relevant person, and the target description text includes the target user's description of the music. For example, the target user can input descriptions such as "Hong Kong and Taiwan songs from the 1990s" or "nostalgic songs."

[0086] Please see Figure 2 In some embodiments, step S102 may include, but is not limited to, steps S201 to S203:

[0087] Step S201: Perform word segmentation on the target description text to obtain multiple target description words;

[0088] Step S202: Extract entity features from the target descriptive words to obtain entity word features;

[0089] Step S203: Based on the preset dictionary and entity word features, filter out the school of thought description words from the target description words.

[0090] In step S201 of some embodiments, when segmenting the target description text, a segmentation tool such as Jieba segmenter can be used to segment the target description text to obtain multiple target description words. Taking Jieba segmenter as an example, firstly, a directed acyclic graph corresponding to the target description text is generated by referring to the dictionary in Jieba segmenter. Then, the shortest path on the directed acyclic graph is found according to the preset selection mode and the dictionary. The target description text is then truncated according to the shortest path, or the target description text is directly truncated to obtain target description words.

[0091] Furthermore, for descriptive words not in the dictionary, Hidden Markov Models (HMMs) can be used for new word discovery. Specifically, the positions B, M, E, and S of a character in the comment text segment are used as hidden states, and the character itself is the observed state. Here, B / M / E / S represent appearing at the beginning, middle, end, and as a single character forming a word, respectively. A dictionary file is used to store the probability matrices, initial probability vectors, and transition probability matrices between characters. The Viterbi algorithm is then used to solve for the maximum possible hidden states, thereby obtaining the target descriptive words.

[0092] In step S202 of some embodiments, named entity algorithms or the like can be used to extract entity features from the target descriptive words to obtain the entity word features corresponding to each target descriptive word.

[0093] In step S203 of some embodiments, the preset dictionary includes multiple reference word features and reference descriptive words. Each reference word feature corresponds to a reference descriptive word. Similarity is calculated between the entity word features and each reference word feature in the preset dictionary to obtain a similarity value between each reference word feature and the entity word feature. The reference word feature with the highest similarity is taken as the corresponding feature for each entity word feature. Based on the correspondence between the reference word features and the reference descriptive words, the reference descriptive word corresponding to each entity word feature is extracted and used as a preliminary descriptive word. Finally, based on the part of speech and meaning of the preliminary descriptive words, preliminary descriptive words representing music genre content are selected as target descriptive words.

[0094] Through the above steps S201 to S203, it is relatively convenient to extract the specific descriptive words of the target object to music from the target description text, and identify the vocabulary content of the music genre in the specific descriptive words, so as to obtain the genre description words. This enables the music that meets the requirements to be retrieved from the candidate music based on the genre description words, thereby improving the accuracy of music retrieval.

[0095] Please see Figure 3In some embodiments, step S103 may include, but is not limited to, steps S301 to S304:

[0096] Step S301: Perform spectral transformation on the candidate music based on a preset function to obtain the Mel cepstral features of the candidate music;

[0097] Step S302: Perform feature segmentation on the Mel cepstral features to obtain multiple candidate spectral segments;

[0098] Step S303: Dimensionality reduction is performed on the candidate spectral segments to obtain the target spectral segments;

[0099] Step S304: The target spectrum segment is spliced ​​to obtain the candidate music spectrum sequence.

[0100] In step S301 of some embodiments, a preset function can be called from a commonly used speech processing database. For example, the speech processing database can be the librosa library, and the preset function can be the librosa.feature.melspectrogram() function in the librosa library. The librosa.feature.melspectrogram() function can be used directly to extract spectral features from candidate music and obtain the Mel cepstral features of the candidate music.

[0101] In step S302 of some embodiments, the Mel-frequency cepstral feature can be further segmented according to a preset time parameter, cutting the Mel-frequency cepstral feature into multiple candidate spectral segments of equal duration. The audio duration and time parameter of each candidate spectral segment are the same. The specific value of the time parameter can be set according to actual needs and is not limited. For example, if the time parameter is 50ms, then the Mel-frequency cepstral feature is cut into multiple candidate spectral segments of equal duration in 50ms units.

[0102] In step S303 of some embodiments, each candidate spectrum segment can be flattened to reduce the two-dimensional features of the candidate spectrum segment to one-dimensional features, thereby obtaining the target spectrum segment. A target spectrum segment is the smallest musical unit token of the candidate music.

[0103] In step S304 of some embodiments, the series of target spectrum segments are merged into a whole according to the time sequence of the target spectrum segments to obtain a candidate music spectrum sequence, that is, the candidate music spectrum sequence is composed of multiple smallest music unit tokens.

[0104] Through the above steps S301 to S304, the audio feature information in the candidate music can be extracted relatively easily, and the audio feature information of the candidate music can be processed into an audio sequence suitable for the model to extract. This can effectively improve the learning ability of the neural network model on the audio features of the candidate music and help improve the accuracy of the model in feature extraction.

[0105] Please see Figure 4 In some embodiments, step S104 may include, but is not limited to, steps S401 to S402:

[0106] Step S401: Based on the preset coding network, feature extraction is performed on the candidate music spectrum sequence to obtain a preliminary music genre representation vector, wherein the coding network includes at least two transformer encoders.

[0107] Step S402: Fine-tune the initial music genre representation vector to obtain the candidate music genre representation vector.

[0108] In step S401 of some embodiments, the aforementioned neural network model includes an encoding network, which includes at least two transformer encoders. Multiple transformer encoders can effectively capture the long-range dependencies between target spectral segments in the input candidate music spectrum sequence. Based on the correlation between the target spectral segments, the overall genre representation information of the candidate music spectrum sequence is extracted, resulting in a preliminary music genre representation vector.

[0109] Since the network structure of each transformer encoder is basically the same, taking the feature extraction process of the first transformer encoder as an example, the candidate music spectrum sequence is input into the first transformer encoder of the encoding network. The encoding part of the transformer encoder first extracts the audio information of the candidate music spectrum sequence, obtaining the first candidate spectral latent vector. This first candidate spectral latent vector includes the basic spectral content information of the candidate music. The encoding part is formed by sequentially connecting an encoding layer, an attention mechanism layer, a normalization layer, and an activation layer. After obtaining the first candidate spectral latent vector, the decoding part of the transformer encoder performs multi-head attention calculation and activation processing on the first candidate spectral latent vector to obtain the first music genre representation vector. The decoding part is formed by sequentially connecting a multi-head attention mechanism layer, a normalization layer, an activation layer, and a fully connected layer. The first music genre representation vector is then input into the second transformer encoder of the encoding network. The above process is repeated based on the second transformer encoder, and so on, until the output of the last transformer encoder is used as the initial music genre representation vector.

[0110] It should be noted that, in order to balance the accuracy of feature extraction and the training efficiency of the encoding network, the transformer encoder in the above encoding network can be set to four, that is, four transformer encoders are connected in sequence. The candidate music spectrum sequence is input into the first transformer encoder for feature extraction, the output of the first transformer encoder is used as the input of the next transformer encoder, and so on, and the output of the last transformer encoder is used as the initial music genre representation vector.

[0111] In step S402 of some embodiments, when fine-tuning the preliminary music genre representation vector, the vector size of the preliminary music genre representation vector can be adjusted based on preset standard size parameters so that the adjusted vector size of the preliminary music genre representation vector meets the actual requirements, thereby obtaining the candidate music genre representation vector.

[0112] Through the above steps S401 to S402, the long-distance dependencies between target spectrum segments in the candidate music spectrum sequence can be captured well, and the correlation between target spectrum segments can be determined based on this long-distance dependency. Thus, the overall genre representation information of the candidate music spectrum sequence can be extracted based on the correlation between target spectrum segments, which can realize the fine-grained representation information of the genre of the input music and help improve the accuracy of music retrieval.

[0113] Please see Figure 5 In some embodiments, step S105 may include, but is not limited to, steps S501 to S503:

[0114] Step S501: Based on the preset candidate genre categories and candidate music genre representation vectors, the candidate music is scored by genre to obtain a music genre score.

[0115] Step S502: Filter the candidate genre categories according to the music genre score to obtain multiple target genre tags for the candidate music;

[0116] Step S503: Obtain genre tag data based on the target genre tag.

[0117] In step S501 of some embodiments, the candidate music genre representation vector can be scored based on a preset genre classifier and candidate genre categories to obtain a music genre score. Specifically, the genre classifier can be a softmax classifier. A probability distribution of the candidate music genre representation vector in each candidate music genre category is created based on the softmax classifier, and the probability distribution vector generated according to the probability distribution is used as the music genre score of the candidate music in the candidate music genre category.

[0118] In step S502 of some embodiments, since the music genre score can intuitively reflect the probability that the candidate music belongs to each candidate genre category, the music genre score is compared with a preset threshold. If the candidate genre category with a music genre score greater than the preset threshold is retained, the retained candidate genre category is used as the target genre label of the candidate music. Therefore, the candidate genre category can be filtered based on the music genre score to obtain multiple target genre labels of the candidate music.

[0119] In step S503 of some embodiments, the music genre score of each target genre tag can be compared, and the target genre tags can be sorted in descending order according to the music genre score to obtain a music genre sequence of candidate music, and this music genre sequence can be used as genre tag data.

[0120] For example, candidate genre categories include pop, rock, electronic, and folk. The genre scores for candidate music in these four categories are 0.5, 0.1, 0.13, and 0.27, respectively. The preset threshold is 0.15. Based on the comparison, pop and folk are selected as the target genre labels for the candidate music. Since 0.5 > 0.27, the target genre labels are sorted in descending order according to the genre score, resulting in the genre label data for the candidate music as [pop, folk].

[0121] Through the above steps S501 to S503, the music genre category of the candidate music can be clearly determined based on the candidate music genre representation vector. It is possible to predict the probability distribution of the candidate music in each candidate music genre category based on the candidate music genre representation vector, which can improve the accuracy of genre identification of candidate music.

[0122] Please see Figure 6 In some embodiments, step S106 includes, but is not limited to, steps S601 to S604:

[0123] Step S601: Analyze the relevance of genre description words and genre tag data to obtain a relevance score;

[0124] Step S602: Based on the relevance score, target tags are selected from the genre tag data, wherein the target tags represent the same music genre as the genre description words.

[0125] Step S603: Obtain the total number of tags for the target tag;

[0126] Step S604: Filter candidate music based on the total number of tags to obtain the target music.

[0127] In step S601 of some embodiments, when scoring the relevance between genre description words and genre tag data, cosine similarity algorithm, Euclidean distance, etc. can be used to calculate the similarity value between genre description words and each genre category in the genre tag data, and the relevance score of the similarity value between the genre description words and each genre category in the genre tag data can be used.

[0128] In step S602 of some embodiments, the relevance score can clearly reflect whether the genre category and the music genre represented by the genre description words are consistent. Therefore, by comparing the relevance score with a preset score threshold, if the relevance score in the genre tag data is greater than the preset score threshold, it indicates that the genre category is consistent with the music genre represented by the genre description words. The genre category in the genre tag data with a relevance score greater than the preset score threshold is taken as the target tag, that is, the target tag represents the same music genre as the genre description words. Therefore, candidate music in the genre tag data that has a target tag representing the same music genre as the genre description words can be directly taken as the target music.

[0129] In step S603 of some embodiments, when there are multiple genre description terms, even if the genre tag data of candidate music contains target tags representing the same music genre as the genre description terms, the candidate music may not necessarily meet the retrieval needs of the target object. Therefore, it is necessary to further filter the candidate music. This filtering can be based on the total number of target tags. When the total number of tags is large, it indicates that the music genre involved in the candidate music is highly similar to the music genre queried by the target object, and is more in line with the needs of the target object. Therefore, the total number of target tags is statistically analyzed using statistical functions such as the SUM function to obtain the total number of tags.

[0130] In step S604 of some embodiments, different tag number thresholds can be set based on different numbers of genre description words. For example, when there is one genre description word, the tag number threshold is 1; when there are three genre description words, the tag number threshold is 2; when there are five genre description words, the tag number threshold is 3, and so on. Therefore, different tag number thresholds are extracted according to the number of genre description words, and the tag number threshold is compared with the total number of tags. When the total number of tags is greater than the tag number threshold, it indicates that the music genre involved in the candidate music is highly similar to the music genre queried by the target object. Therefore, candidate music with a total number of tags greater than the tag number threshold can be used as target music.

[0131] Through the above steps S601 to S604, it is relatively convenient to filter out target music from multiple candidate music based on genre description words, so as to select the genre category that better meets the search needs of the target object, which can effectively reduce the amount of search data and improve the accuracy of music search.

[0132] Please see Figure 7 In some embodiments, step S107 may include, but is not limited to, steps S701 to S702:

[0133] Step S701: Sort the target music based on the total number of tags of the target music to obtain a music search list;

[0134] Step S702: The music search list is sent back to the target object.

[0135] In step S701 of some embodiments, since the total number of tags can reflect the degree of matching between the target music and the genre content of the target description text input by the target object, the more tags there are, the higher the degree of matching between the target music and the genre content of the target description text input by the target object. Therefore, the total number of tags for each target music can be compared, and based on the total number of tags for different target music, this series of target music can be sorted in descending order, with the target music with more tags in the first position and the target music with fewer tags in the last position, to obtain a music search list.

[0136] In step S702 of some embodiments, the music search list is directly fed back to the target user, or the target music that is at the top of the music search list is selected and fed back to the target object.

[0137] Furthermore, when providing feedback on target music, various formats such as tables, tree diagrams, or blocks can be used to present it to the target audience. For example, based on the primary genre of the target music, the music can be divided into multiple blocks, and different genre blocks can be colored and displayed to the target audience.

[0138] By using the above steps S701 and S702, we can improve the accuracy of music retrieval, reduce communication costs, and increase the diversity of target music display.

[0139] The music retrieval method of this application embodiment obtains target description text and candidate music. The target description text includes the target object's description of the music. Word recognition is performed on the target description text to obtain genre description terms, which can easily identify the vocabulary describing music genres in the target description text and obtain genre description terms, enabling the retrieval of music that meets the requirements from the candidate music based on the genre description terms. Furthermore, spectral transformation is performed on the candidate music to obtain candidate music spectrum sequences, which can easily extract audio feature information from the candidate music. Furthermore, based on the candidate music spectrum sequences, the corresponding candidate music genre representation vector is obtained, which enables fine-grained representation of the genre information of the candidate music, thus improving the accuracy of music retrieval. Furthermore, the candidate music genre is identified based on the candidate music genre representation vector to obtain the genre tag data of the candidate music. The candidate music is then filtered based on the genre description words and genre tag data to obtain the target music. The target music is then fed back to the target object. This method can conveniently filter the target music from multiple candidate music based on genre description words to select the target music whose genre category better matches the target object's search needs. It can effectively reduce the amount of search data and improve the accuracy of music search.

[0140] Please see Figure 8This application also provides a music retrieval device that can implement the above-described music retrieval method. The device includes:

[0141] The acquisition module 801 is used to acquire the target description text and candidate music, wherein the target description text includes the description content of the target object on the music;

[0142] The word recognition module 802 is used to perform word recognition on the target description text to obtain the words describing the genre;

[0143] The spectrum transformation module 803 is used to perform spectrum transformation on the candidate music to obtain the candidate music spectrum sequence;

[0144] The genre representation extraction module 804 is used to obtain the candidate music genre representation vector corresponding to the candidate music based on the candidate music spectrum sequence.

[0145] The genre identification module 805 is used to identify the genre of candidate music based on the candidate music genre representation vector to obtain the genre label data of the candidate music.

[0146] The music filtering module 806 is used to filter candidate music based on genre description words and genre tag data to obtain target music;

[0147] The result feedback module 807 is used to provide the target music back to the target object.

[0148] The specific implementation of this music retrieval device is basically the same as the specific implementation of the music retrieval method described above, and will not be repeated here.

[0149] This application also provides an electronic device, which includes: a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for communication between the processor and the memory. When the program is executed by the processor, it implements the aforementioned music retrieval method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0150] Please see Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:

[0151] The processor 901 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0152] The memory 902 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called and executed by the processor 901 using the music retrieval method of the embodiments of this application.

[0153] The input / output interface 903 is used to implement information input and output;

[0154] The communication interface 904 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0155] Bus 905 transmits information between various components of the device (e.g., processor 901, memory 902, input / output interface 903, and communication interface 904);

[0156] The processor 901, memory 902, input / output interface 903, and communication interface 904 are connected to each other within the device via bus 905.

[0157] This application also provides a computer-readable storage medium storing one or more programs that can be executed by one or more processors to implement the music retrieval method described above.

[0158] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0159] The music retrieval method, music retrieval device, electronic device, and computer-readable storage medium provided in this application embodiment acquire target descriptive text and candidate music. The target descriptive text includes the target object's description of the music. Word recognition is performed on the target descriptive text to obtain genre description terms, which can easily identify the vocabulary describing music genres in the target descriptive text, thus obtaining genre description terms. This allows for the retrieval of music that meets the requirements from candidate music based on the genre description terms. Furthermore, spectral transformation is performed on the candidate music to obtain candidate music spectrum sequences, which can easily extract audio feature information from the candidate music. Further, based on the candidate music spectrum sequences, the corresponding candidate music genre representation vector is obtained, enabling fine-grained representation of the genre information of the candidate music, which is beneficial to improving the accuracy of music retrieval. Furthermore, the candidate music genre is identified based on the candidate music genre representation vector to obtain the genre tag data of the candidate music. The candidate music is then filtered based on the genre description words and genre tag data to obtain the target music. The target music is then fed back to the target object. This method can conveniently filter the target music from multiple candidate music based on genre description words to select the target music whose genre category better matches the target object's search needs. It can effectively reduce the amount of search data and improve the accuracy of music search.

[0160] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0161] It will be understood by those skilled in the art that Figure 1-7 The technical solutions shown do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0162] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0163] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0164] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0165] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0166] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0167] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0168] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0169] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0170] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A music retrieval method, characterized in that, The method includes: Obtain target description text and candidate music, wherein the target description text includes the target object's description of the music; Word recognition is performed on the target descriptive text to obtain genre description words; Perform spectral transformation on the candidate music to obtain a candidate music spectral sequence; Based on the candidate music spectrum sequence, obtain the candidate music genre representation vector corresponding to the candidate music; Based on the candidate music genre representation vector, the candidate music is identified into a genre to obtain the genre tag data of the candidate music. The candidate music is filtered based on the genre description words and genre tag data to obtain the target music; The target music is fed back to the target object; The step of obtaining the candidate music genre representation vector corresponding to the candidate music based on the candidate music spectrum sequence includes: Based on a pre-set encoding network containing at least two transformer encoders, the long-distance dependencies between target spectrum segments in the candidate music spectrum sequence are captured. The overall genre representation information of the candidate music spectrum sequence is extracted based on the correlation between the target spectrum segments to obtain a preliminary music genre representation vector. The vector size of the preliminary music genre representation vector is adjusted based on preset standard size parameters to obtain the candidate music genre representation vector.

2. The music retrieval method according to claim 1, characterized in that, The word recognition process for the target descriptive text, yielding genre description words, includes: The target description text is segmented to obtain multiple target description words; Entity features are extracted from the target descriptive words to obtain entity word features; Based on a preset dictionary and the features of the entity words, the descriptive words of the school of thought are selected from the target descriptive words.

3. The music retrieval method according to claim 1, characterized in that, The process of filtering candidate music based on the genre description words and genre tag data to obtain target music includes: A relevance score is obtained by performing a relevance assessment on the descriptive terms of the genre and the tag data of the genre; Based on the relevance score, target tags are selected from the genre tag data, wherein the target tags and the genre description words represent the same music genre; Obtain the total number of tags for the target tag; The candidate music is filtered based on the total number of tags to obtain the target music.

4. The music retrieval method according to claim 3, characterized in that, The step of feeding back the target music to the target object includes: The target music is sorted based on the total number of tags in the target music to obtain a music search list; The music search list is then fed back to the target object.

5. The music retrieval method according to claim 1, characterized in that, The step of performing spectral transformation on the candidate music to obtain a candidate music spectral sequence includes: The candidate music is subjected to spectral transformation based on a preset function to obtain the Mel cepstral features of the candidate music; The Mel cepstral features are segmented to obtain multiple candidate spectral segments; The candidate spectral segment is dimensionality reduced to obtain the target spectral segment; The target spectrum segment is spliced ​​together to obtain the candidate music spectrum sequence.

6. The music retrieval method according to any one of claims 1 to 5, characterized in that, The process of identifying the genre of candidate music based on the candidate music genre representation vector to obtain genre label data for the candidate music includes: The candidate music is scored based on the preset candidate genre categories and the candidate music genre representation vectors to obtain a music genre score. The candidate genre categories are filtered based on the music genre scores to obtain multiple target genre tags for the candidate music. Based on the target genre tag, the genre tag data is obtained.

7. A music retrieval device, characterized in that, The device includes: The acquisition module is used to acquire target description text and candidate music, wherein the target description text includes the description content of the music by the target object; The word recognition module is used to perform word recognition on the target description text to obtain genre description words; The spectrum transformation module is used to perform spectrum transformation on the candidate music to obtain a candidate music spectrum sequence; The genre representation extraction module is used to obtain the candidate music genre representation vector corresponding to the candidate music based on the candidate music spectrum sequence. The genre identification module is used to identify the genre of the candidate music based on the candidate music genre representation vector, and obtain the genre tag data of the candidate music. The music filtering module is used to filter the candidate music based on the genre description words and the genre tag data to obtain the target music; The result feedback module is used to feed back the target music to the target object; The step of obtaining the candidate music genre representation vector corresponding to the candidate music based on the candidate music spectrum sequence includes: Based on a pre-set encoding network containing at least two transformer encoders, the long-distance dependencies between target spectrum segments in the candidate music spectrum sequence are captured. The overall genre representation information of the candidate music spectrum sequence is extracted based on the correlation between the target spectrum segments to obtain a preliminary music genre representation vector. The vector size of the preliminary music genre representation vector is adjusted based on preset standard size parameters to obtain the candidate music genre representation vector.

8. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the music retrieval method as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the music retrieval method according to any one of claims 1 to 6.