Music recommendation method and device, and XR device
Patent Information
- Application Number
- CN202511242483.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2045-09-02
AI Technical Summary
[0004]本发明提供一种音乐推荐方法、装置及XR设备,用以解决现有音乐推荐结果的准确性较差的问题
[0015] The music recommendation method, apparatus, and XR device provided by this invention, in response to a user-inputted music recommendation command, acquire a first scene text, convert the first scene text into a first semantic vector to capture deep semantic relationships, and extract a first keyword from the first scene text to extract core information. Then, based on the first semantic vector and the first keyword, search results are retrieved from a preset music index library to determine the user's target recommended music. Through the collaborative retrieval of semantic vectors and keywords, it exhibits strong adaptability to scene texts of different styles and complexities, making recommended music more aligned with the emotional tone and functional needs of the user's current situation, thereby improving the accuracy of recommendation results and user satisfaction. Simultaneously, it can quickly narrow the search scope and accurately locate the target music, improving search response speed.
Smart Images

Figure CN120744169B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a music recommendation method, apparatus, and XR device. Background Technology
[0002] In recent years, scene-aware intelligent content recommendation, such as automatically matching or generating suitable music based on the user's context, has become an important direction for improving user experience. However, existing technical solutions typically employ a single-keyword retrieval recommendation model, which can only achieve a literal match, resulting in poor accuracy of subsequent music recommendation results.
[0003] Therefore, improving the accuracy of music recommendation results is a technical problem that urgently needs to be solved. Summary of the Invention
[0004] This invention provides a music recommendation method, apparatus, and XR device to solve the problem of poor accuracy in existing music recommendation results.
[0005] This invention provides a music recommendation method, comprising: In response to the user's input of a music recommendation command, obtain the first scene text; The first scene text is converted into a first semantic vector, and the first keyword in the first scene text is extracted; Based on the first semantic vector and the first keyword, the search results are obtained from the preset music index library; Based on the search results, the user's target recommended music is determined.
[0006] According to a music recommendation method provided by the present invention, the step of retrieving search results from a preset music index based on the first semantic vector and the first keyword includes: Based on the first semantic vector, the first candidate recommended music and the first similarity are retrieved from the preset music index library; Based on the first keyword, a second candidate recommended music and a second similarity are retrieved from the preset music index library; The first similarity and the second similarity are weighted and summed to obtain a fusion similarity; wherein, the search result includes the first candidate recommended music, the second candidate recommended music, and the fusion similarity.
[0007] According to a music recommendation method provided by the present invention, the preset music index library is constructed as follows: Obtain a music file and extract a first music feature from the music file, wherein the first music feature includes at least one of an audio fingerprint and music metadata; Based on the first music feature, the pre-constructed mapping knowledge base between the second scene text and the second music feature, the scene tag text corresponding to the music file is obtained; Convert the scene label text into a second semantic vector; The preset music index library is constructed based on the music file, the first music feature, the scene tag text, and the second semantic vector.
[0008] According to a music recommendation method provided by the present invention, the retrieval result includes candidate recommended music and fusion similarity, and the step of determining the user's target recommended music based on the retrieval result includes: When at least one value in the fusion similarity is detected to be greater than or equal to a preset similarity threshold, the user's target recommended music is determined from the candidate recommended music based on the fusion similarity. When the detected fusion similarity is less than the preset similarity threshold, a music generation task is generated based on the first semantic vector and the first keyword. The music generation task is then submitted to the music generation model through an asynchronous queue to generate the target recommended music.
[0009] According to a music recommendation method provided by the present invention, before determining the user's target recommended music from the candidate recommended music based on the fusion similarity, the method includes: Obtain callback logs within a preset time period, the callback logs including historical recommended music; The step of determining the user's target recommended music from the candidate recommended music based on the fusion similarity includes: Based on the fusion similarity and the historical recommended music, the user's target recommended music is selected from the candidate recommended music.
[0010] According to a music recommendation method provided by the present invention, the music generation task includes a preset callback function. After generating the music generation task based on the first semantic vector and the first keyword, and submitting the music generation task to the music generation model through an asynchronous queue to generate target recommended music, the method further includes: The third musical feature of the target recommended music is extracted through the preset callback function; The preset music index is updated based on the target recommended music, the third music feature, the first keyword, and the first semantic vector.
[0011] According to a music recommendation method provided by the present invention, the step of converting the first scene text into a first semantic vector includes: The first scene text is input into the text embedding model for vector conversion, and the first semantic vector of a preset dimension is obtained by the text embedding model.
[0012] According to a music recommendation method provided by the present invention, after determining the user's target recommended music based on the search results, the method further includes: Play the target recommended music and monitor user interaction data in real time; Based on the user interaction data, adjust the retrieval weight parameters of the preset music index library.
[0013] The present invention also provides a music recommendation device, comprising: The acquisition module is used to acquire the first scene text in response to the user's input music recommendation command; The processing module is used to convert the first scene text into a first semantic vector and extract the first keyword from the first scene text; The retrieval module is used to retrieve retrieval results from a preset music index based on the first semantic vector and the first keyword; The determination module is used to determine the user's target recommended music based on the search results.
[0014] The present invention also provides an XR device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the music recommendation method as described above.
[0015] The music recommendation method, apparatus, and XR device provided by this invention, in response to a user-inputted music recommendation command, acquire a first scene text, convert the first scene text into a first semantic vector to capture deep semantic relationships, and extract a first keyword from the first scene text to extract core information. Then, based on the first semantic vector and the first keyword, search results are retrieved from a preset music index library to determine the user's target recommended music. Through the collaborative retrieval of semantic vectors and keywords, it exhibits strong adaptability to scene texts of different styles and complexities, making recommended music more aligned with the emotional tone and functional needs of the user's current situation, thereby improving the accuracy of recommendation results and user satisfaction. Simultaneously, it can quickly narrow the search scope and accurately locate the target music, improving search response speed. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0017] Figure 1 This is a system architecture diagram of the music recommendation system provided by the present invention; Figure 2 This is one of the flowcharts illustrating the music recommendation method provided by the present invention; Figure 3 This is the second flowchart illustrating the music recommendation method provided by the present invention; Figure 4 This is the third flowchart illustrating the music recommendation method provided by the present invention; Figure 5 This is a schematic diagram of the music recommendation device provided by the present invention; Figure 6 This is a schematic diagram of the structure of the XR device provided by the present invention. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0019] This invention proposes a music recommendation method, apparatus, and XR device, which are described below in conjunction with... Figures 1-6 Describe it.
[0020] Figure 1 This is a system architecture diagram of the music recommendation system provided by the present invention, such as... Figure 1 As shown, the music recommendation system includes an Extended Reality (XR) device 01 and a server 02.
[0021] XR device 01 includes, but is not limited to, VR (Virtual Reality) devices, AR (Augmented Reality) devices, and MR (Mixed Reality) devices, used to execute the music recommendation method of this invention. Server 02 can be a server used to train text embedding models and build music index libraries.
[0022] After training the text embedding model and building the music index library, server 02 deploys the trained text embedding model and music index library to XR device 01. XR device 01 can convert scene text into semantic vectors based on the text embedding model and perform retrieval based on the music index library to obtain retrieval results.
[0023] In other embodiments of the present invention, XR device 01 may not deploy a text embedding model and a music index library. Instead, server 02 determines the target recommended music based on the text embedding model and the music index library and sends it to XR device 01. Specifically, XR device 01 obtains first scene text and sends it to server 02. Server 02 converts the received first scene text into a first semantic vector and extracts the first keyword from the first scene text. Then, based on the first semantic vector and the first keyword, it retrieves search results from the music index library and determines the target recommended music based on the search results. Server 02 then sends the target recommended music to XR device 01 for playback.
[0024] Figure 2 This is one of the flowcharts illustrating the music recommendation method provided by the present invention, such as... Figure 2 As shown, the music recommendation method includes steps S110, S120, S130 and S140.
[0025] Step S110: In response to the user's input music recommendation command, obtain the first scene text.
[0026] In this embodiment, the music recommendation method is applied to XR devices, but it can also be applied to terminal devices and servers such as smartphones, tablets, laptops, and desktop computers.
[0027] XR devices refer to wearable or portable devices that integrate virtual and real environments through hardware and software technologies to enable human-computer interaction. XR devices include, but are not limited to, VR devices, AR devices, and MR devices. VR devices use computer technology to simulate and generate a three-dimensional virtual space, allowing users to immerse themselves and interact with it, gaining a truly immersive experience. AR devices use technology to merge virtual information with the real world, overlaying it onto real-world scenes in real time to enhance sensory experience. MR devices mix the real and virtual worlds to create a new visual environment that simultaneously contains physical entities and virtual information, allowing users to interact with these physical entities and virtual information in real time.
[0028] This music recommendation method is suitable for scenarios where music recommendations are based on scene images or scene text.
[0029] Music recommendation commands can be triggered in ways including but not limited to: voice triggering, gesture triggering, and virtual interface operation. For example, a voice triggering scenario could be: the user says "Please recommend some music"; a gesture triggering scenario could be: the user makes a specific music recommendation gesture; and a virtual interface operation could be: a floating control panel is generated, and the user selects a virtual music recommendation button using a gamepad / gesture.
[0030] In response to the user's input music recommendation command, the corresponding scene text is retrieved. Scene text is text that describes a scene. For example, "autumn dusk" or "coffee shop". To distinguish it from the scene text used later to build the preset music index library, the scene text retrieved here is denoted as the first scene text.
[0031] One method of acquisition is to capture scene images using an XR device, and then convert the scene images to obtain the first scene text. During the scene image conversion, Vision-Language Models (VLMs) can be used. The scene image is input into the visual language model for image-to-text conversion, resulting in the scene text output by the visual language model.
[0032] As another method of acquisition, scene images can be captured using the sensors of an XR device, and the user's physiological state information can be obtained. This physiological state information includes, but is not limited to, heart rate, eye gaze direction, head orientation, and gait parameters. Combining the scene images and physiological state information, a first scene text can be generated. For example, if green grass is identified from the scene image, and the user's behavior is determined to be a walk based on heart rate and gait parameters, the scene text "outdoor walk" can be generated.
[0033] As another acquisition method, users can input scene speech through the microphone of the XR device. Scene speech refers to the voice data describing the current scene. Upon receiving the user's input scene speech, the system recognizes it to obtain the first scene text. During scene speech recognition, an ASR (Automatic Speech Recognition) model can be used. The scene speech is input into the ASR model for speech recognition, and the scene text output by the ASR model is obtained.
[0034] As another way to obtain the information, users can input scene text through external devices (such as user terminals or Bluetooth keyboards) or the virtual keyboard of XR devices. In this case, the first scene text input by the user can be directly obtained.
[0035] Step S120: Convert the first scene text into a first semantic vector and extract the first keyword from the first scene text.
[0036] The first scene text is converted into a semantic vector, denoted as the first semantic vector.
[0037] When performing semantic vector transformation, the first scene text can be input into a text embedding model for vector transformation, resulting in the first semantic vector output by the text embedding model. The text embedding model can include, but is not limited to: BGE (BAAI General Embedding), GTE (General Text Embeddings), M3E (Moka Massive Mixed Embedding), BERT (Bidirectional Encoder Representations from Transformers), etc.
[0038] At the same time, extract the keywords from the first scene text and record them as the first keyword.
[0039] When extracting keywords, you can use the TextRank (graph-based ranking) algorithm, TF-IDF (Term Frequency-Inverse Document Frequency) algorithm, domain keyword list or part-of-speech rule extraction, or you can use a deep learning model (such as the BERT model) for extraction.
[0040] Step S130: Based on the first semantic vector and the first keyword, retrieve the search results from the preset music index library.
[0041] The construction process of the preset music index library is as follows: Obtain music files; extract the first music feature of the music files, which includes at least one of audio fingerprint and music metadata; match the scene tag text corresponding to the music files based on the first music feature and a pre-built mapping knowledge base between the second scene text and the second music feature; convert the scene tag text into a second semantic vector; and construct the preset music index library based on the music files, the first music feature, the scene tag text, and the second semantic vector. The specific execution process can be found in the following embodiments, which will not be elaborated here.
[0042] As one implementation method, the process of obtaining the search results is as follows: based on the first semantic vector, a first candidate recommended music and a first similarity are retrieved from a preset music index library; simultaneously, based on the first keyword, a second candidate recommended music and a second similarity are retrieved from the preset music index library; the search results include the first candidate recommended music, the first similarity, the second candidate recommended music, and the second similarity.
[0043] As another implementation method, the process of obtaining the search results is as follows: based on the first semantic vector, a first candidate recommended music and a first similarity are retrieved from a preset music index library; simultaneously, based on the first keyword, a second candidate recommended music and a second similarity are retrieved from the preset music index library; the first similarity and the second similarity are weighted and summed to obtain a fused similarity; wherein, the search results include the first candidate recommended music, the second candidate recommended music, and the fused similarity. The specific execution process can be referred to in the following embodiments, which will not be elaborated here.
[0044] Step S140: Based on the search results, determine the user's target recommended music.
[0045] After obtaining the search results, the user's target recommended music is determined based on the search results.
[0046] When the search results include the first candidate recommended music, the first similarity, the second candidate recommended music, and the second similarity, the first candidate recommended music and the second candidate recommended music can be sorted according to the first similarity and the second similarity. The candidate recommended music with the highest similarity is selected as the target recommended music.
[0047] When the search results include a first candidate recommended music, a second candidate recommended music, and a fusion similarity score, one implementation method is to sort the first and second candidate recommended music based on the fusion similarity score, and select the top n candidate recommended music scores as the target recommended music. Another implementation method involves detecting whether all fusion similarities are greater than or equal to a preset similarity threshold. If at least one fusion similarity value is detected to be greater than or equal to the preset similarity threshold, then the candidate recommended music with a fusion similarity greater than or equal to the preset similarity threshold is selected as the target recommended music. If all fusion similarities are detected to be less than the preset similarity threshold, then a music generation task is generated based on the first semantic vector and the first keyword, and the music generation task is submitted to the music generation model via an asynchronous queue to generate the target recommended music.
[0048] The music recommendation method provided in this invention responds to a user-inputted music recommendation command by acquiring a first scene text, converting the first scene text into a first semantic vector to capture deep semantic relationships, and extracting a first keyword from the first scene text to extract core information. Then, based on the first semantic vector and the first keyword, it retrieves search results from a preset music index database to determine the user's target recommended music. Through the collaborative retrieval of semantic vectors and keywords, it exhibits strong adaptability to scene texts of different styles and complexities, making the recommended music more aligned with the emotional tone and functional needs of the user's current situation, thereby improving the accuracy of the recommendation results and user satisfaction. Simultaneously, it can quickly narrow the search scope and accurately locate the target music, improving search response speed.
[0049] Based on any of the above embodiments Figure 3 This is the second flowchart of the music recommendation method provided by the present invention, as shown below. Figure 3 As shown, step S130 includes: step S131, step S132 and step S133.
[0050] Step S131: Based on the first semantic vector, retrieve the first candidate recommended music and the first similarity from the preset music index library.
[0051] Call the Elasticsearch vector retrieval plugin to load vector data from the second semantic vector field in the preset music index. The similarity between the first semantic vector and the second semantic vector corresponding to each music file in the preset music index is calculated and denoted as the first similarity. This similarity can be represented by cosine similarity. Music with a first similarity ≥ a first preset value (e.g., 0.7) is selected as candidate recommended music and denoted as the first candidate recommended music, and the corresponding first similarity is retained. For example, the first semantic vector of "gym training" entered by the user and the second semantic vector of music tagged "sports and fitness scene" (such as "Hot Blood") in the preset music index have a first similarity of 0.85, and the second semantic vector of music tagged "outdoor running scene" (such as "Running") have a first similarity of 0.78, both of which are included in the first candidate recommended music.
[0052] Step S132: Based on the first keyword, retrieve the second candidate recommended music and the second similarity from the preset music index library.
[0053] The full-text search functionality of Elasticsearch is invoked. An inverted index is built based on the scene tag text and music metadata fields (such as style and rhythm) in the first music feature. The BM25 (Best Matching 25) algorithm is used to calculate the keyword matching score, which is denoted as the second similarity. Music with a second similarity ≥ a second preset value (e.g., 0.6) is selected as candidate recommended music, denoted as the second candidate recommended music, and the corresponding second similarity is retained. For example, the first keywords "running" and "rhythm" match music in the preset music index with the tag "sports and fitness scene" and music metadata "rhythm: fast" (such as "Running") with a second similarity of 0.82, and match music with the tag "outdoor cycling scene" (such as "Sprint") with a second similarity of 0.75; both are included in the second candidate recommended music.
[0054] It should be noted that the first preset value and the second preset value can be the same or different.
[0055] Step S133: The first similarity and the second similarity are weighted and summed to obtain a fusion similarity; wherein the retrieval result includes the first candidate recommended music, the second candidate recommended music and the fusion similarity.
[0056] The fusion similarity is obtained by weighting and summing the first similarity, second similarity, first preset weight coefficient, and second preset weight coefficient. The specific formula is: Fusion Similarity = First Similarity × First Preset Weight Coefficient + Second Similarity × Second Preset Weight Coefficient. The first and second preset weight coefficients can be pre-set static values or dynamically adjusted values.
[0057] For example, assuming the first preset weight coefficient is 0.7 and the second preset weight coefficient is 0.3, in the above example, for the music "Hot Blood", its first similarity is 0.85 and its second similarity is 0.70, so its fusion similarity can be calculated as 0.85×0.7+0.70×0.3=0.805; for the music "Running", its first similarity is 0.78 and its second similarity is 0.82, so its fusion similarity can be calculated as 0.78×0.7+0.82×0.3=0.792; for the music "Sprint", its first similarity is 0 (not reaching the first preset value, calculated as 0) and its second similarity is 0.75, so its fusion similarity can be calculated as 0×0.7+0.75×0.3=0.225.
[0058] The music recommendation method provided in this invention employs a hybrid search and retrieval (HSR) approach. Vector search provides semantic understanding and similarity matching, while keyword search ensures accurate matching and interpretability. By fusing multi-dimensional retrieval using semantic vectors and keywords, the limitations of single retrieval methods can be overcome, improving the overall retrieval effect and thus enhancing the accuracy and scenario adaptability of music recommendations.
[0059] Based on any of the above embodiments Figure 4 This is the third flowchart of the music recommendation method provided by the present invention, as shown below. Figure 4 As shown, the construction method of the preset music index library includes: steps S10, S20, S30 and S40.
[0060] S10, Obtain a music file and extract a first music feature from the music file, wherein the first music feature includes at least one of an audio fingerprint and music metadata.
[0061] Music file acquisition methods include, but are not limited to: 1) Local acquisition: reading music files located in a preset storage path on a local storage device; 2) Network acquisition: accessing a music resource server and downloading music files; 3) Generating music files through a music generation model. The music generation model may include, but is not limited to: Suno V3 (the third-generation music generation model developed by Suno Corporation), Udio, and MuseNet (an AI music creation system based on deep neural networks developed by OpenAI).
[0062] A music fingerprint is a unique, compact digital identifier obtained by converting an audio signal. Audio fingerprints of music files can be generated using the Chromaprint algorithm (a core component of the open-source project AcoustID).
[0063] Music metadata includes, but is not limited to, song title, artist, album, duration, genre, tempo, and bitrate. It can be obtained by parsing the ID3 tag (metadata standard embedded in audio files) and file header information. For untagged music files, the music metadata is obtained by comparing its audio features with those in an online music database.
[0064] S20, based on the first music feature, the pre-constructed mapping knowledge base of the second scene text and the second music feature, the scene tag text corresponding to the music file is obtained by matching.
[0065] A pre-built knowledge base mapping second scene text to second music features, including the mapping relationship between scene text and music features. The mapping knowledge base can be built using a MySQL database to store the mapping relationship between second scene text and second music features.
[0066] For example, the second scene text is "coffee shop leisure scene", and the corresponding second music features are: audio fingerprint (low frequency energy accounts for 30%-40%, spectrum entropy value 0.6-0.7) and music metadata (style is "jazz" or "light music", rhythm is "soothing", duration is 180-300 seconds).
[0067] For example, the second scene text is "sports and fitness scene", and the corresponding second music features are: audio fingerprint (high frequency energy ratio of 40%-50%, spectrum entropy value of 0.8-0.9) and music metadata (style of "rock" and "electronic", rhythm of "fast", duration of 120-200 seconds).
[0068] The first music feature is compared with the second music feature in the mapping knowledge base. A weighted cosine similarity algorithm (e.g., audio fingerprint weight is 0.6, music metadata weight is 0.4) is used to calculate the similarity, which is recorded as the third similarity. The second scene text with a third similarity ≥ a third preset value (e.g., 85%) is selected as the scene label text corresponding to the music file. For example, the rock song "Hot Blood" is matched with the scene label text "sports and fitness scene".
[0069] S30, convert the scene label text into a second semantic vector.
[0070] The scene label text is converted into a semantic vector, denoted as the second semantic vector. The semantic vector conversion method for the scene label text is the same as that for the first scene text.
[0071] When performing semantic vector transformation, the scene label text can be input into a text embedding model for vector transformation, resulting in a second semantic vector output by the text embedding model. The text embedding model can include, but is not limited to, BGE models, GTE models, M3E models, BERT models, etc.
[0072] S40, construct the preset music index library based on the music file, the first music feature, the scene tag text, and the second semantic vector.
[0073] The default music index is built using the Elasticsearch search engine and includes: (1) Index structure design: It may include, but is not limited to, 6 fields, namely, the storage path of the music file, the first music feature, the scene tag text, the second semantic vector, the index weight, and the index creation time. (2) Index building process: The above information of the music files is formatted according to the index structure and written to the index database in batches through Elasticsearch's Bulk API. For music metadata and scene text tags in music features, an inverted index mechanism is used to establish a mapping relationship between keywords and documents. A vector retrieval plugin is enabled for the second semantic vector field to support vector nearest neighbor queries based on cosine similarity. Audio fingerprints can be stored separately through a hash index.
[0074] Furthermore, the preset music index storage capacity supports dynamic expansion and contraction, meaning it can automatically adjust the number of nodes based on real-time load. When the load exceeds the fourth preset value (e.g., 80%), new nodes can be automatically added (e.g., expanding from 5 nodes to 6) to alleviate the pressure on the existing nodes, preventing retrieval delays or failures due to insufficient resources and ensuring that retrieval requests always respond quickly. When the load remains below the fifth preset value (e.g., 20%) for an extended period, the number of nodes can be automatically reduced to minimize hardware resource waste.
[0075] Furthermore, the default music index is split into multiple shards, each storing a portion of the data and distributed across different nodes. When a new node is added, shard migration can be used to transfer some shards from the original node to the new node, achieving load balancing. Shard migration is set to complete within 10 minutes to ensure that the retrieval service is not significantly affected during the migration process, maintaining high availability of the index.
[0076] The music recommendation method provided in this invention can achieve efficient processing and index construction of music files through the above-described manner, and can support subsequent fast retrieval based on keywords or semantic vectors.
[0077] Furthermore, prior to step S40, the following steps are also included: The second semantic vector is subjected to vector quantization to obtain the quantized second semantic vector.
[0078] At this point, step S40 includes: The preset music index library is constructed based on the music file, the first music feature, the scene tag text, and the quantized second semantic vector.
[0079] Vector quantization is a lossy data compression technique used to optimize storage and retrieval efficiency. Vector quantization methods include, but are not limited to, product quantization (PQ), binary quantization (BQ), and scalar quantization (SQ).
[0080] Furthermore, product quantization can be used in this embodiment, which is particularly suitable for high-dimensional vector retrieval, large-scale datasets, and scenarios with high requirements for storage and retrieval efficiency.
[0081] In this embodiment, by employing vector quantization technology to compress the second semantic vector, the data volume can be significantly reduced, effectively saving storage space, improving transmission efficiency, and lowering costs. Simultaneously, this technology can significantly optimize subsequent retrieval performance. Experiments show that compared to the uncompressed original data, the retrieval speed using this solution is increased by up to 5 times, and the concurrent retrieval capacity of a single node can reach 500+.
[0082] Based on any of the above embodiments, the retrieval results include candidate recommended music and fusion similarity, and step S140 includes: step S141 and step S142.
[0083] Step S141: When it is detected that at least one value in the fusion similarity is greater than or equal to a preset similarity threshold, the user's target recommended music is determined from the candidate recommended music based on the fusion similarity.
[0084] The search results include candidate recommended music and fusion similarity, where candidate recommended music includes one or more pieces, and correspondingly, fusion similarity includes one or more pieces.
[0085] After obtaining the search results, it is checked whether the fusion similarity is greater than or equal to the preset similarity threshold (e.g., 0.85). If at least one value in the fusion similarity is greater than or equal to the preset similarity threshold, it is determined that there is music that is very suitable for the text in the first scenario. At this time, the candidate recommended music with a fusion similarity greater than or equal to the preset similarity threshold is selected as the target recommended music for the user.
[0086] It should be noted that if there are multiple candidate recommended music tracks with a fusion similarity greater than or equal to the preset similarity threshold, they can be recommended and played in descending order of fusion similarity.
[0087] Step S142: When the detected fusion similarity is less than the preset similarity threshold, a music generation task is generated based on the first semantic vector and the first keyword. The music generation task is submitted to the music generation model through an asynchronous queue to generate target recommended music.
[0088] If the detected similarity scores are all less than the preset similarity threshold, it is determined that there is no suitable music in the text of the first scenario. In this case, the music generation model is used to generate the target recommended music.
[0089] Specifically, a music generation task can be generated first based on the first semantic vector and the first keyword. This music generation task, in addition to the first semantic vector and the first keyword, can also include a task ID (identity document), task generation time, and a preset callback function. The music generation task can be encapsulated in JSON (JavaScript Object Notation) format.
[0090] Then, the music generation task is submitted to the music generation model through an asynchronous queue, so that the music generation model can generate target recommended music based on the music generation task.
[0091] Asynchronous queues can be implemented using Async-Queue (a tool for managing the execution order and concurrency control of asynchronous tasks). Compared to traditional Kafka queues (distributed message middleware), while Kafka supports asynchronous task distribution, its callback mechanism lacks flexibility, making it difficult to synchronize the generated results (target recommended music) to the preset music index in real time, resulting in a 1-2 second end-to-end process delay. Async-Queue, however, can successfully compress the processing latency of the generated results to within 500ms, achieving a single cluster throughput of 2000+ QPS (concurrently generated tasks).
[0092] Furthermore, upon receiving a music generation task, the asynchronous queue can assign a priority label to the task, allowing submissions to be ordered by priority. The priority labeling method can employ a numerical hierarchy, where numbers represent priorities; for example, smaller numbers indicate higher priority, such as P0 (highest), P1, P2, and P3 (lowest).
[0093] Furthermore, the marking rules may include: 1) Static predefined: that is, a task type-priority mapping table is pre-established; 2) Dynamic rules: based on the static predefined priorities, the task priorities are automatically adjusted by detecting specific conditions at runtime to achieve intelligent scheduling.
[0094] For example, real-time scene tasks can be labeled as P1, and batch generation tasks can be labeled as P3.
[0095] Priority labeling allows for the processing of higher-priority tasks, thereby maximizing resource utilization while ensuring system responsiveness.
[0096] Music generation models can be based on Suno V3, Udio, or MuseNet.
[0097] Furthermore, the music generation model can be fine-tuned, and then the target recommended music can be generated through the fine-tuned music generation model.
[0098] Fine-tuning refers to further training an existing music generation model using a specific multitrack music dataset of millions of tracks to adjust its internal parameters (including multitrack generation parameters) and improve its performance on multitrack music generation tasks. Multitrack music refers to music composed of multiple independent tracks, such as drum tracks, bass tracks, piano tracks, vocal tracks, and string tracks. Each track can be edited independently (volume, panning, effects), and then mixed into a complete piece. Fine-tuning can significantly improve the performance of music generation models on multitrack music generation tasks, specifically optimizing track separation, time synchronization accuracy, musical harmony, and overall mixing quality.
[0099] The music recommendation method provided in this invention flexibly switches between "filtering existing resources" and "generating new resources" modes by judging the relationship between fusion similarity and a threshold. This fully utilizes existing music resources to improve recommendation efficiency while also creating new music resources through a music generation model when necessary, thus combining the advantages of resource reuse and innovative generation, making the recommendation method more intelligent and flexible. Furthermore, using an asynchronous queue to submit music generation tasks avoids blocking the main music recommendation process, ensuring that the system can still efficiently respond to other user operations while processing music generation tasks. This balances resource allocation between real-time recommendation and complex music generation tasks, improving the overall smoothness and stability of the system.
[0100] Based on any of the above embodiments, before step S141, the method further includes: Obtain callback logs within a preset time period, including historical music recommendations.
[0101] In this embodiment, considering that users would have a poor experience if they consistently received similar music recommendations over a period of time, the system first retrieves callback logs for a preset time period when at least one value in the fusion similarity is detected to be greater than or equal to a preset similarity threshold, indicating that there is music highly suitable for the text in the first scenario. The preset time period can be pre-set, such as one day or one week. The callback logs include historical music recommendations, and may also include user ID, playback timestamp, and callback time.
[0102] At this point, step S141 includes: Based on the fusion similarity and the historical recommended music, the user's target recommended music is selected from the candidate recommended music.
[0103] After obtaining the historical recommended music, the user's target recommended music is selected from the candidate recommended music based on the fusion similarity and the historical recommended music.
[0104] Specifically, music with a similarity greater than or equal to a preset similarity threshold can be selected from the candidate recommended music, and then historical recommended music can be filtered out to obtain the target recommended music.
[0105] The music recommendation method provided in this invention obtains music that has been recommended within a preset time period (i.e., historical recommended music), and then further filters out historical recommended music based on the fusion similarity of candidate recommended music, thereby ensuring the freshness of the target recommended music and improving the user experience.
[0106] Based on any of the above embodiments, the music generation task includes a preset callback function, and after step S142, it further includes steps S150 and S160.
[0107] Step S150: Extract the third music feature of the target recommended music through the preset callback function.
[0108] In this embodiment, the music generation task includes a preset callback function. This preset callback function is used to precisely trigger the extraction of third-party music features and the updating of the preset music index library after the target recommended music is generated.
[0109] Specifically, after the music generation model generates the target recommended music, a preset callback function is triggered. First, the music features of the target recommended music are extracted, denoted as the third music feature. The third music feature includes at least one of the following: audio fingerprint and music metadata. The music fingerprint is a unique, compact numerical identifier obtained by converting an audio signal. The audio fingerprint of a music file can be generated using the Chromaprint algorithm (a core component of the open-source project AcoustID). Music metadata includes, but is not limited to: song title, artist, album, duration, genre, tempo, and bitrate. It can be obtained by parsing the ID3 tag (metadata standard embedded in audio files) and file header information of the music file. For untagged music files, music metadata is obtained by comparing the audio features with those in an online music database.
[0110] Step S160: Update the preset music index library based on the target recommended music, the third music feature, the first keyword, and the first semantic vector.
[0111] Then, based on the target recommended music, the third music feature, the first keyword, and the first semantic vector, the preset music index is updated.
[0112] Specifically, a new document is generated according to the index structure, which is the same as before and may include, but is not limited to, six fields: the storage path of the music file, the first music feature, the scene tag text, the second semantic vector, the index weight, and the index creation time. The scene tag text for the target recommended music can be determined based on the first keyword. Alternatively, it can be determined based on the first scene text. Then, the new document is written to the pre-defined music index repository via the Elasticsearch Index API, following the index building process described above. Furthermore, an optimistic locking mechanism can be used during the write process to avoid concurrency conflicts.
[0113] Furthermore, before updating the preset music index, the similarity of existing music in the preset music index can be compared using audio fingerprints. If the similarity is less than a fifth preset value (e.g., 95%), then an update is performed. If the similarity is greater than or equal to the fifth preset value, it is determined to be a duplicate, and no update is performed. The music recommendation method provided in this invention decouples asynchronous generation from subsequent processes through callback functions, ensuring that feature extraction and index updates are triggered only after the target recommended music is generated. This reduces system resource blocking, improves the reliability of asynchronous task processing, and is suitable for dynamic music generation and index maintenance in high-concurrency scenarios.
[0114] Furthermore, when music generation fails, the error parameters and error type can be recorded through a callback function and the task can be retried, automatically adjusting the music generation parameters, such as reducing the BPM (Beat Per Minute).
[0115] Error parameters refer to the specific set of input parameter values (including the BPM value used at that time) that triggered a generation failure. The error callback function is responsible for capturing and recording these parameters, providing a basis for subsequent automatic parameter adjustments (such as reducing the BPM), thereby realizing an intelligent retry mechanism. Error parameters may include, but are not limited to: BPM, number of tracks, generation duration, etc.
[0116] Error types may include, but are not limited to: invalid input, resource limit exceeded, audio rendering failure, etc.
[0117] In this embodiment, automatically optimizing generation conditions through failure feedback increases the probability of the music generation task ultimately succeeding. This mechanism improves the system's robustness, especially when faced with resource constraints or unreasonable parameter settings, by adaptively adjusting parameters to attempt to complete the music generation task.
[0118] Based on any of the above embodiments, the step of "converting the first scene text into a first semantic vector" includes: The first scene text is input into the text embedding model for vector conversion, and the first semantic vector of a preset dimension is obtained by the text embedding model.
[0119] In this embodiment, the process of obtaining the first semantic vector is as follows: inputting the first scene text into the text embedding model for vector conversion, and obtaining the first semantic vector of preset dimensions output by the text embedding model.
[0120] The text embedding model is a BGE model with a default dimension of 1024. Furthermore, a BGE-large version can be selected to support the output of 1024-dimensional semantic vectors.
[0121] Compared to existing text embedding models such as GTE, M3E, and BERT, the BGE model demonstrates significant advantages in core scenarios such as semantic retrieval and text clustering, especially in its stronger semantic decomposition capabilities in dynamic scenarios, effectively avoiding information loss. Furthermore, the use of 1024 dimensions improves accuracy in semantically complex scenarios requiring fine-grained matching. For example, low-dimensional vectors might confuse "romantic rainy night" and "sad rainy night" into similar vectors, while the 1024-dimensional high-dimensional vector can distinguish the detailed features of "romantic" (lights, encounters) and "sad" (loneliness, partings) through different dimensions, making it more suitable for scenarios with complex text semantics (such as artistic descriptions and interwoven emotions) and where fine-grained matching of musical features is required.
[0122] Existing text embedding models have limited ability to extract semantics from complex scenes; for example, the accuracy of semantic vector similarity calculation for "autumn dusk + coffee shop + jazz music" is less than 70%. However, the accuracy of the first semantic vector generated by the BGE-large model, which has 1024 dimensions, can be improved to over 96%.
[0123] The music recommendation method provided in this invention converts the first scene text into a first semantic vector of a preset dimension through a text embedding model. This transforms the abstract semantics of the scene text into a structured vector representation, breaking through the limitations of traditional text processing that relies solely on literal keywords. It can effectively capture deeper information such as synonyms, contextual relationships, and implicit intentions, providing a unified vector space for subsequent operations such as accurate matching and weighted fusion, and expanding the application potential of the system in multimodal interaction scenarios.
[0124] Based on any of the above embodiments, after step S140, the method further includes steps S170 and S180.
[0125] Step S170: Play the target recommended music and monitor user interaction data in real time.
[0126] After determining the user's target music recommendation, play the target music and monitor user interaction data in real time.
[0127] It should be understood that the target recommended music may include one or more songs. If multiple songs are included, they can be played in order of recommendation (sorted from largest to smallest according to fusion similarity).
[0128] User interaction data includes both positive and negative interactions. Positive interactions include, but are not limited to, liking, favoriteing, repeating, and sharing music; negative interactions include, but are not limited to, skipping the current song and adding the song to the "dislikes" list.
[0129] Step S180: Adjust the retrieval weight parameters of the preset music index library based on the user interaction data.
[0130] Then, based on the monitored user interaction data, the retrieval weight parameters of the preset music index library are adjusted.
[0131] Specifically, adjustments can be made when user interaction occurs, i.e., when user interaction data is detected. Adjustments can also be made periodically, such as every 2 hours or every day. In this method, user interaction data can be stored in a callback log first, and then the callback log can be retrieved periodically to adjust the retrieval weight parameters of the preset music index library.
[0132] During adjustments, the monitored user interaction data can be encapsulated into structured data and passed to a preset music index library via another pre-defined callback function. For example, user interaction data, such as music ID, user ID, behavior type, and playback duration, can be encapsulated into the following structured data: {"music_id":"123","user_id":"u456","behavior":"skip","play_time":10}.
[0133] Search weight parameters are parameters used during the query process to adjust the degree of influence of different semantic vectors or keywords on the final relevance score of the music file. For example, in a positive interaction, the search weight of the corresponding music file's semantic vector can be increased by 0.05; in a negative interaction, the search weight of the corresponding music file's semantic vector can be decreased by 0.1. The adjustment rules for search weight parameters can be preset and are not specifically limited here.
[0134] The music recommendation method provided in this invention monitors user interaction data to adjust the retrieval weight parameters of a preset music index, enabling music recommendations to respond to changes in user preferences, improving recommendation accuracy, and increasing user satisfaction with recommended music.
[0135] The following example illustrates the application scenario of the music recommendation method provided by this invention in XR devices.
[0136] Users take images of the current scene using XR devices (such as wearing XR glasses) and convert the captured images into text descriptions (i.e., scene text), such as "cyberpunk style + street night scene".
[0137] Then, the scene text is converted into a 1024-dimensional semantic vector, and keywords are extracted from the scene text. Based on the semantic vector and keywords, a hybrid search and retrieval is performed in a pre-built music index library using the Elasticsearch search engine to obtain existing music with a fusion similarity ≥ 0.85. The obtained existing music is filtered through the 24-hour played music in the callback log to ensure that the recommendations are not duplicated before playback. If no music is found, an Async-Queue music generation task is triggered and a callback is awaited. After the target recommended music is generated, the callback function automatically triggers feature extraction and writes the extracted music features, their corresponding music files, semantic vectors, and scene tag texts into the pre-built music index library.
[0138] When playing target recommended music, monitor user interaction data and dynamically adjust the retrieval weight parameters of the preset music index library based on callback signals of user interaction data (such as likes and comments).
[0139] The music recommendation device provided by the present invention will be described below. The music recommendation device described below can be referred to in correspondence with the music recommendation method described above.
[0140] Figure 5 This is a schematic diagram of the music recommendation device provided by the present invention, as shown below. Figure 5 As shown, the device includes an acquisition module 510, a processing module 520, a retrieval module 530, and a determination module 540; wherein: The acquisition module 510 is used to acquire the first scene text in response to the user's input music recommendation command; The processing module 520 is used to convert the first scene text into a first semantic vector and extract the first keyword from the first scene text; The retrieval module 530 is used to retrieve retrieval results from a preset music index based on the first semantic vector and the first keyword; The determination module 540 is used to determine the user's target recommended music based on the search results.
[0141] The music recommendation device provided in this invention responds to a user-inputted music recommendation command by acquiring a first scene text, converting the first scene text into a first semantic vector to capture deep semantic relationships, and extracting a first keyword from the first scene text to extract core information. Then, based on the first semantic vector and the first keyword, it retrieves search results from a preset music index database to determine the user's target recommended music. Through the collaborative retrieval of semantic vectors and keywords, it exhibits strong adaptability to scene texts of different styles and complexities, making the recommended music more aligned with the emotional tone and functional needs of the user's current situation, thereby improving the accuracy of the recommendation results and user satisfaction. Simultaneously, it can quickly narrow the search scope and accurately locate the target music, improving search response speed.
[0142] According to a music recommendation device provided by the present invention, the retrieval module 530 is specifically used for: Based on the first semantic vector, the first candidate recommended music and the first similarity are retrieved from the preset music index library; Based on the first keyword, a second candidate recommended music and a second similarity are retrieved from the preset music index library; The first similarity and the second similarity are weighted and summed to obtain a fusion similarity; wherein, the search result includes the first candidate recommended music, the second candidate recommended music, and the fusion similarity.
[0143] According to a music recommendation device provided by the present invention, the preset music index library is constructed in the following manner: Obtain a music file and extract a first music feature from the music file, wherein the first music feature includes at least one of an audio fingerprint and music metadata; Based on the first music feature, the pre-constructed mapping knowledge base between the second scene text and the second music feature, the scene tag text corresponding to the music file is obtained; Convert the scene label text into a second semantic vector; The preset music index library is constructed based on the music file, the first music feature, the scene tag text, and the second semantic vector.
[0144] According to a music recommendation device provided by the present invention, the search results include candidate recommended music and fusion similarity, and the determining module 540 is specifically used for: When at least one value in the fusion similarity is detected to be greater than or equal to a preset similarity threshold, the user's target recommended music is determined from the candidate recommended music based on the fusion similarity. When the detected fusion similarity is less than the preset similarity threshold, a music generation task is generated based on the first semantic vector and the first keyword. The music generation task is then submitted to the music generation model through an asynchronous queue to generate the target recommended music.
[0145] According to a music recommendation device provided by the present invention, the determining module 540 is further specifically used for: Obtain callback logs within a preset time period, the callback logs including historical recommended music; Based on the fusion similarity and the historical recommended music, the user's target recommended music is selected from the candidate recommended music.
[0146] According to a music recommendation device provided by the present invention, the music generation task includes a preset callback function, and the music recommendation device further includes an update module for: The third musical feature of the target recommended music is extracted through the preset callback function; The preset music index is updated based on the target recommended music, the third music feature, the first keyword, and the first semantic vector.
[0147] According to the present invention, a music recommendation device, a processing module 520, is specifically used for: The first scene text is input into the text embedding model for vector conversion, and the first semantic vector of a preset dimension is obtained by the text embedding model.
[0148] According to a music recommendation device provided by the present invention, the music recommendation device further includes an adjustment module for: Play the target recommended music and monitor user interaction data in real time; Based on the user interaction data, adjust the retrieval weight parameters of the preset music index library.
[0149] It should be noted that the music recommendation device provided in this embodiment of the invention can implement all the method steps implemented in the above-mentioned music recommendation method embodiment and can achieve the same technical effect. Therefore, the parts and beneficial effects that are the same as those in the method embodiment will not be described in detail here.
[0150] Figure 6 An example is a schematic diagram of the physical structure of an XR device, such as... Figure 6As shown, the XR device may include a processor 610, a communications interface 620, a memory 630, and a communication bus 640, wherein the processor 610, communications interface 620, and memory 630 communicate with each other via the communication bus 640. The processor 610 can invoke logical instructions in the memory 630 to execute a music recommendation method, which includes: in response to a user-inputted music recommendation instruction, obtaining first scene text; converting the first scene text into a first semantic vector and extracting a first keyword from the first scene text; retrieving search results from a preset music index based on the first semantic vector and the first keyword; and determining the user's target recommended music based on the search results.
[0151] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0152] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0153] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A music recommendation method characterized by, include: In response to the user's input of a music recommendation command, obtain the first scene text; The first scene text is converted into a first semantic vector of 1024 dimensions by a BGE model, and a first keyword in the first scene text is extracted; wherein the BGE model selects BGE large version to support the output of the 1024-dimensional semantic vector. Based on the first semantic vector and the first keyword, search results are obtained from a preset music index library; the preset music index library is built based on the Elasticsearch search engine, and the index structure includes a second semantic vector field for vector retrieval and an inverted index for keyword retrieval, the inverted index being built based on scene tag text and music metadata fields in the first music feature; the search results include candidate recommended music and fusion similarity; When the detected fusion similarity is less than a preset similarity threshold, a music generation task is generated based on the first semantic vector and the first keyword. The music generation task is submitted to the music generation model through an asynchronous queue (Async-Queue) to generate target recommended music. The music generation task includes a preset callback function. When the target recommended music is successfully generated, the third music feature of the target recommended music is extracted through the preset callback function; the third music feature includes at least one of audio fingerprint and music metadata; Based on the target recommended music, the third music feature, the first keyword, and the first semantic vector, a new document is generated according to the index structure. The new document is written to the preset music index via the Elasticsearch Index API to update the preset music index. Before updating the preset music index, the similarity of existing music in the preset music index is compared using the audio fingerprint. If the similarity is less than a fifth preset value, an update is performed; if the similarity is greater than or equal to the fifth preset value, it is determined to be a duplicate, and no update is performed. When the target recommended music generation fails, the error parameters and error type are recorded through the preset callback function, and the music generation parameters are adjusted to trigger a retry of the music generation task. The step of retrieving search results from a preset music index based on the first semantic vector and the first keyword includes: retrieving a first candidate recommended music and a first similarity from the preset music index based on the first semantic vector; retrieving a second candidate recommended music and a second similarity from the preset music index based on the first keyword; and performing a weighted summation of the first similarity and the second similarity to obtain a fusion similarity; wherein the search results include the first candidate recommended music, the second candidate recommended music, and the fusion similarity; Specifically, the full-text search function of Elasticsearch is invoked, and an inverted index is constructed based on the scene tag text and the music metadata field in the first music feature. The keyword matching score is calculated using the BM25 algorithm and recorded as the second similarity. Music with the second similarity greater than or equal to the second preset value is selected as candidate recommended music and recorded as the second candidate recommended music, and the corresponding second similarity is retained.
2. The music recommendation method according to claim 1, characterized in that, The preset music index library is constructed as follows: Obtain a music file and extract a first music feature from the music file, wherein the first music feature includes at least one of an audio fingerprint and music metadata; Based on the first music feature, the pre-constructed mapping knowledge base between the second scene text and the second music feature, the scene tag text corresponding to the music file is obtained; Convert the scene label text into a second semantic vector; The preset music index library is constructed based on the music file, the first music feature, the scene tag text, and the second semantic vector.
3. The music recommendation method according to claim 1, characterized in that, After obtaining the search results from the preset music index based on the first semantic vector and the first keyword, the process further includes: When at least one value in the fusion similarity is detected to be greater than or equal to a preset similarity threshold, the user's target recommended music is determined from the candidate recommended music based on the fusion similarity.
4. The music recommendation method according to claim 3, characterized in that, Before determining the user's target recommended music from the candidate recommended music based on the fusion similarity, the process includes: Obtain callback logs within a preset time period, the callback logs including historical recommended music; The step of determining the user's target recommended music from the candidate recommended music based on the fusion similarity includes: Based on the fusion similarity and the historical recommended music, the user's target recommended music is selected from the candidate recommended music.
5. The music recommendation method according to claim 4, characterized in that, After determining the user's target recommended music, the process also includes: Play the target recommended music and monitor user interaction data in real time; Based on the user interaction data, adjust the retrieval weight parameters of the preset music index library.
6. A music recommendation device, characterized in that, include: The acquisition module is used to acquire the first scene text in response to the user's input music recommendation command; The processing module is used to convert the first scene text into a 1024-dimensional first semantic vector using a BGE model, and to extract the first keyword from the first scene text; wherein, the BGE model is selected as BGE. Large version to support the output of 1024-dimensional semantic vectors; The retrieval module is used to retrieve retrieval results from a preset music index library based on the first semantic vector and the first keyword; the preset music index library is built based on the Elasticsearch search engine, and the index structure includes a second semantic vector field for vector retrieval and an inverted index for keyword retrieval, the inverted index being built based on scene tag text and music metadata fields in the first music feature; the retrieval results include candidate recommended music and fusion similarity; The determining module is configured to generate a music generation task based on the first semantic vector and the first keyword when the detected fusion similarity is less than a preset similarity threshold, and submit the music generation task to the music generation model through an asynchronous queue (Async-Queue) to generate target recommended music; wherein, the music generation task includes a preset callback function; When the target recommended music is successfully generated, the third music feature of the target recommended music is extracted through the preset callback function; the third music feature includes at least one of audio fingerprint and music metadata; Based on the target recommended music, the third music feature, the first keyword, and the first semantic vector, a new document is generated according to the index structure. The new document is written to the preset music index via the Elasticsearch Index API to update the preset music index. Before updating the preset music index, the similarity of the audio fingerprint to the existing music in the preset music index is compared. If the similarity is less than a fifth preset value, the update is performed. If the similarity is greater than or equal to the fifth preset value, it is determined to be a duplicate and no update is performed. When the target recommended music generation fails, the error parameters and error type are recorded through the preset callback function, and the music generation parameters are adjusted to trigger a retry of the music generation task; The step of retrieving search results from a preset music index based on the first semantic vector and the first keyword includes: retrieving a first candidate recommended music and a first similarity from the preset music index based on the first semantic vector; retrieving a second candidate recommended music and a second similarity from the preset music index based on the first keyword; and performing a weighted summation of the first similarity and the second similarity to obtain a fusion similarity; wherein the search results include the first candidate recommended music, the second candidate recommended music, and the fusion similarity; Specifically, the full-text search function of Elasticsearch is invoked, and an inverted index is constructed based on the scene tag text and the music metadata field in the first music feature. The keyword matching score is calculated using the BM25 algorithm and recorded as the second similarity. Music with the second similarity greater than or equal to the second preset value is selected as candidate recommended music and recorded as the second candidate recommended music, and the corresponding second similarity is retained.
7. An XR device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the music recommendation method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Song list pushing method and device, computer equipment and storage medium
CN111078931A
Music searching method and device, equipment and storage medium
CN116680437A
Online content recommendation method and system based on semantic discovery
CN120492737A