Intelligent photo album system supporting voice-on-demand music
By collecting and converting user voice information into text in the intelligent album system, and using a large language model for cross-modal data mapping and retrieval, the problem of the inability to accurately retrieve music in the existing technology is solved, and the system's intelligence and interactive performance are improved.
Patent Information
- Application Number
- CN202510602550.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-08-19
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing smart photo album system cannot accurately retrieve the music files required by the user through large-scale analysis based on the voice command information entered by the user, and the intelligent interactive use is insufficient.
User voice information is collected through the intelligent album terminal device and converted into text, and feature extraction and search is used for large language model servers. Combined with keyword search, semantic matching and graph reasoning, the language-multimodal comparison model maps text and audio cross-modal data to a unified semantic space for precise search. Finally, the retrieved music is played by the music platform server.
It realizes high-precision audio retrieval and playback of user voice input, improving the intelligence level of the system and human-computer interaction performance.
Smart Images

Figure CN120508676A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of smart photo albums, and in particular relates to a smart photo album system that supports voice-on-demand music. Background Art
[0002] With the popularity and development of smartphones, people are increasingly using their phones to take photos, recording the daily lives of themselves and their families. However, these photos often only exist in the phone's storage space, rarely being viewed or shared, and their value and significance cannot be fully reflected. Therefore, smart photo albums came into being.
[0003] For example, the invention disclosed in announcement number CN118409675A discloses an AI multimedia circular scrolling display smart touch family album, which relates to the technical field of smart family albums. It is a large-screen device that can store and display mobile phone photos, has a multi-screen scrolling function, a voice interaction function with a large language model, and a multimodal image recognition, retrieval, classification, and control function based on artificial intelligence. It can easily transfer photos on the mobile phone to the large-screen device, and through flexible configuration of the split screen to divide into multiple small screens, scroll and play the photos you want to play, and can realize the function of arranging and retrieving photos on the display screen through voice interaction, thereby increasing the viewing and interactivity of the photos, and improving the family atmosphere and emotional communication.
[0004] However, the above technical solution cannot accurately retrieve the music files required by the user through large-scale model analysis based on the voice command information input by the user. It does not support voice-on-demand music and its interactive use is not intelligent enough. Therefore, we propose an intelligent photo album system that supports voice-on-demand music. Summary of the Invention
[0005] The purpose of the present invention is to provide an intelligent photo album system that supports voice-on-demand music, so as to solve the problem that the existing technology proposed in the above background technology cannot accurately retrieve the music files required by the user through large-scale model analysis based on the voice command information input by the user, and it does not support voice-on-demand music and the interactive use is not intelligent enough.
[0006] To achieve the above-mentioned object, the present invention provides the following technical solutions: an intelligent photo album system supporting voice-on-demand music, comprising an intelligent photo album terminal device, an intelligent photo album server, and a large language model server;
[0007] The smart album terminal device is interactively connected to the smart album server or the music platform server, the smart album server is interactively connected to the large language model server or the music platform server, and the large language model server is interactively connected to the music platform server;
[0008] The smart photo album terminal device is used to collect user language information and convert the voice information into dialogue text;
[0009] The large language model server is used to extract features from the conversation text and transmit them to the music platform server for retrieval.
[0010] Preferably, the smart album terminal device includes a microphone and a smart album APP module, the microphone is interactively connected to the smart album APP module or the music platform server, and the smart album APP module is also interactively connected to the speaker and the display or the music platform server.
[0011] Preferably, the speaker and display screen are both arranged in the smart album terminal, the speaker is used to play the detected music, and the display screen is used to display the dialogue text generated by user voice conversion, as well as to play the retrieved music MV and display the lyrics.
[0012] Preferably, the microphone is used to collect voice information input by the user.
[0013] Preferably, the smart album server is further provided with a smart question-answering service module, and the smart question-answering service module is used to conduct interactive question-answering with the user.
[0014] Preferably, a large language model system module is provided in the large language model server, and the large language model system module is interactively connected with the intelligent question and answer service module.
[0015] Preferably, the large language model system module constructs a vector database, combines keyword search, semantic matching and graph reasoning, and uses a language-multimodal contrast model to map text and audio cross-modal data into a unified semantic space for accurate audio retrieval.
[0016] Preferably, a music platform system module is provided in the music platform server, and the music platform system module is interactively connected with the large language model system module or the smart album APP module.
[0017] Compared with the prior art, the present invention has the following beneficial effects:
[0018] The present invention can convert the information input by the user's voice into text information, analyze the text information through a large model, accurately retrieve the audio, and play the retrieved audio, which greatly improves the audio retrieval accuracy, has a high overall intelligence level, and has good human-computer interaction performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is a schematic diagram of the overall structure of the existing system;
[0020] Figure 2This is a schematic diagram of the system framework structure of Example 1 of the present invention;
[0021] Figure 3 This is a schematic diagram of the system framework structure of Example 2 of the present invention;
[0022] Figure 4 This is a schematic diagram of the system framework structure of Example 3 of the present invention;
[0023] Figure 5 This is a schematic diagram of the system framework structure of Example 4 of the present invention. DETAILED DESCRIPTION
[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0025] Example 1
[0026] See also Figure 2 ,The present invention provides a technical solution: an intelligent photo album system supporting voice-on-demand music, comprising an intelligent photo album terminal device, an intelligent photo album server and a large language model server;
[0027] The smart album terminal device is interactively connected to the smart album server, the smart album server is interactively connected to the large language model server, and the large language model server is interactively connected to the music platform server;
[0028] The smart photo album terminal device is used to collect user language information and convert the voice information into dialogue text;
[0029] The large language model server is used to extract features from the conversation text and transmit them to the music platform server for retrieval.
[0030] The smart album terminal device includes a microphone and a smart album APP module. The microphone is interactively connected to the smart album APP module. The smart album APP module is also interactively connected to the speaker and the display. The microphone is used to collect voice information input by the user.
[0031] A music platform system module is set in the music platform server, and the music platform system module is interactively connected with the large language model system module.
[0032] Furthermore, a speaker and a display screen are both provided in the smart album terminal. The speaker is used to play the detected music, and the display screen is used to display the dialogue text generated by the user's voice conversion, as well as to play the retrieved music MV and display the lyrics.
[0033] Furthermore, an intelligent question-answering service module is also provided in the intelligent photo album server, and the intelligent question-answering service module is used to conduct interactive question-answering with users.
[0034] Furthermore, a large language model system module is set in the large language model server, and the large language model system module is interactively connected with the intelligent question and answer service module.
[0035] Furthermore, the large language model system module builds a vector database, combines keyword search, semantic matching and graph reasoning, and uses a language-multimodal comparison model to map text and audio cross-modal data into a unified semantic space for accurate audio retrieval.
[0036] Specifically,
[0037] The overall architecture of the large language model system module is as follows:
[0038] 1. Build a vector database: Generate semantic vectors by parsing domain documents (such as music metadata and enterprise knowledge base) to support similarity retrieval.
[0039] 2. Knowledge graph fusion: Extract entity relationships to build a graph database to enhance the logical relevance of search results (such as the reasoning link of "singer A → song style → song name").
[0040] 3. Adopt a hybrid retrieval strategy: combine keyword search, semantic matching, and graph reasoning to achieve multi-dimensional data recall.
[0041] 4. Dynamic intent analysis module: Through natural language processing technology, user queries (such as "find a piece of light music suitable for travel") are broken down into multi-label combinations such as scene and style.
[0042] 5. Use language-multimodal contrast models (such as CLaMP) to map cross-modal data such as text and audio into a unified semantic space to improve cross-type retrieval accuracy (such as matching melodies with lyrics).
[0043] 6. Inject industry knowledge bases (such as music genre systems) into the retrieval stage and solve long-tail query problems (such as "Japanese City Pop style songs from the 1980s") through graph completion algorithms.
[0044] 7. Based on the analysis of user click behavior and semantic association, the vector encoder parameters are continuously updated to achieve online learning of the retrieval model.
[0045] Language-Multimodal Contrast Model Construction Method
[0046] 1. Core Architecture Design
[0047] 1. Joint training of multimodal encoders
[0048] A dual-tower structure is adopted: independent text encoders (such as BERT, RoBERTa) and cross-modal encoders (such as CLIP visual encoder) are constructed respectively, and semantic space alignment is achieved through contrastive learning.
[0049] Modal encoder selection:
[0050] Vision: ViT or ResNet extracts image block features and generates visual semantic vectors through self-attention mechanism;
[0051] Auditory: Use Wav2Vec or MFCC to extract audio spectrum features and combine them with time series modeling to generate acoustic representations.
[0052] 2. Projection alignment module
[0053] Constructing a cross-modal projection layer: Mapping the embedding vectors of different modalities into a unified semantic space. For example, by using linear transformations or lightweight Transformers to achieve dimensionality matching between visual / auditory features and text features. 58
[0054] Introducing a learnable adapter: adding an adaptation layer to the output of the pre-trained encoder to preserve the original parameters while enhancing modality alignment capabilities.
[0055] 2. Contrastive Learning Mechanism
[0056] 1. Loss function design
[0057] Adopting InfoNCE loss: maximize the similarity of positive sample pairs (such as image-description matching pairs) and minimize the similarity of negative sample pairs;
[0058] Expanded multi-granularity comparison: Simultaneously aligns global features (entire image) and local features (key object areas) to improve fine-grained semantic capture capabilities.
[0059] 2. Data Enhancement Strategy
[0060] Implement inter-modal enhancement: generate multimodal variants of the same semantic object (e.g. crop and rotate an image and simultaneously generate corresponding text description perturbations);
[0061] Use hybrid negative sampling: combine random negative samples within the batch with manually constructed difficult negative samples (such as image-text pairs with similar scenes but different semantics).
[0062] 3. Training Process Optimization
[0063] 1. Two-stage training mechanism
[0064] Pre-training phase: training comparison targets on large-scale image / text / audio paired data (such as COCO and AudioSet) to learn basic cross-modal alignment capabilities;
[0065] Fine-tuning phase: Add a task head (classifier / generator) to specific task data (such as music-lyrics matching, medical imaging report generation) and perform end-to-end optimization.
[0066] 2. Dynamic resolution processing
[0067] Implement block encoding for high-resolution input: split the image into multiple patches and fuse local and global information through cross-attention (such as the dual encoder mechanism of CogAgent);
[0068] Adaptive scaling: Dynamically adjust input resolution based on device computing power, balancing computational efficiency and feature quality.
[0069] The specific system operation method of this embodiment is as follows:
[0070] Step 1: The microphone collects the voice information input by the user and transmits the voice information to the smart album APP module;
[0071] Step 2: The smart photo album APP module converts the voice information into text information and transmits the text information to the display screen, which displays the text information;
[0072] Step 3: The smart photo album APP module transmits the text information to the smart question-answering service module, and interactive questions and answers can be conducted through the smart question-answering service module;
[0073] Step 4: The smart album server transmits the text information to the large speech model system module. By building a vector database, combining keyword search, semantic matching, and graph reasoning, and using a language-multimodal comparison model, it maps the text and audio cross-modal data into a unified semantic space for accurate audio retrieval.
[0074] Step 5: The large language model system module transmits the text information of the user's music request to the music platform system module and performs retrieval;
[0075] Step 6: The music platform server transmits the retrieved music information to the large language model server;
[0076] Step 7: The large language model server transmits the retrieved music to the smart album server;
[0077] Step 8: The smart album server transmits the retrieved music to the smart album terminal device;
[0078] Step 9: Play the retrieved music MV through the display screen, display the lyrics, and play the retrieved music through the speaker.
[0079] Example 2
[0080] See also Figure 3 , an intelligent photo album system supporting voice-on-demand music, including an intelligent photo album terminal device and an intelligent photo album server;
[0081] The smart album terminal device is interactively connected to the smart album server, and the smart album server is interactively connected to the music platform server; the smart album terminal device is used to collect user language information and convert the voice information into dialogue text.
[0082] The smart album terminal device includes a microphone and a smart album APP module. The microphone is interactively connected to the smart album APP module or the music platform server. The smart album APP module is also interactively connected to the speaker and the display.
[0083] The speaker and display are both set in the smart album terminal. The speaker is used to play the detected music, the display is used to display the dialogue text generated by user voice conversion, as well as play the retrieved music MV and display lyrics. The microphone is used to collect voice information input by the user; the smart album server also has an intelligent question and answer service module, which is used to interact with users for questions and answers; the music platform system module is set in the music platform server, and the music platform system module is interactively connected with the smart album APP module.
[0084] The specific system operation method of this embodiment is as follows:
[0085] Step 1: The microphone collects the voice information input by the user and transmits the voice information to the smart album APP module;
[0086] Step 2: The smart photo album APP module converts the voice information into text information and transmits the text information to the display screen, which displays the text information;
[0087] Step 3: The smart photo album APP module transmits the text information to the smart question-answering service module, and interactive questions and answers can be conducted through the smart question-answering service module;
[0088] Step 4: The intelligent question-answering service module transmits the text information of the user's music request to the music platform system module and performs retrieval;
[0089] Step 5: The music platform server transmits the retrieved music information to the smart album server;
[0090] Step 6: The smart album server transmits the retrieved music to the smart album terminal device;
[0091] Step 7: Play the retrieved music MV through the display screen, display the lyrics, and play the retrieved music through the speaker.
[0092] Example 3
[0093] See also Figure 4 , an intelligent photo album system that supports voice on-demand music, including an intelligent photo album terminal device; the intelligent photo album terminal device is interactively connected to the music platform server, the intelligent photo album terminal device is used to collect user language information and convert the voice information into dialogue text, the intelligent photo album terminal device includes a microphone and an intelligent photo album APP module, the microphone is interactively connected to the intelligent photo album APP module, the intelligent photo album APP module is also interactively connected to the speaker and the display and the music platform server, the speaker and the display are both set in the intelligent photo album terminal, the speaker is used to play the detected music, the display is used to display the dialogue text generated by the user voice conversion, as well as to play the retrieved music MV and display lyrics, and the microphone is used to collect the voice information input by the user; a music platform system module is set in the music platform server, and the music platform system module is interactively connected to the large language model system module or the intelligent photo album APP module.
[0094] The specific system operation method of this embodiment is as follows:
[0095] Step 1: The microphone collects the voice information input by the user and transmits the voice information to the smart album APP module;
[0096] Step 2: The smart photo album APP module converts the voice information into text information and transmits the text information to the display screen, which displays the text information;
[0097] Step 3: The smart album APP module transmits the text information of the user's music request to the music platform system module and performs retrieval;
[0098] Step 4: The music platform server transmits the retrieved music information to the smart album terminal device;
[0099] Step 5: Play the retrieved music MV through the display screen, display the lyrics, and play the retrieved music through the speaker.
[0100] Example 4
[0101] See also Figure 5, an intelligent photo album system that supports voice-on-demand music, including an intelligent photo album terminal device, the intelligent photo album terminal device is used to collect user language information and convert the voice information into dialogue text, the intelligent photo album terminal device includes a microphone and an intelligent photo album APP module, the microphone is interactively connected with the intelligent photo album APP module, the intelligent photo album APP module is also interactively connected with the display screen and the music APP, and the music APP is interactively connected with the speaker and the display screen.
[0102] The specific system operation method of this embodiment is as follows:
[0103] Step 1: The microphone collects the voice information input by the user and transmits the voice information to the smart album APP module;
[0104] Step 2: The smart photo album APP module converts the voice information into text information and transmits the text information to the display screen, which displays the text information;
[0105] Step 3: After the smart photo album app interprets the user's instructions, it calls the music app to play the specified music;
[0106] Step 4: Play the retrieved music MV through the display screen, display the lyrics, and play the retrieved sound through the speaker.
[0107] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. An intelligent photo album system supporting voice-on-demand music, characterized by: Including smart album terminal devices, smart album servers and large language model servers; The smart album terminal device is interactively connected to the smart album server or the music platform server, the smart album server is interactively connected to the large language model server or the music platform server, and the large language model server is interactively connected to the music platform server; The smart photo album terminal device is used to collect user language information and convert the voice information into dialogue text; The large language model server is used to extract features from the conversation text and transmit them to the music platform server for retrieval.
2. The intelligent photo album system supporting voice-on-demand music according to claim 1, characterized in that: The smart album terminal device includes a microphone and a smart album APP module. The microphone is interactively connected to the smart album APP module, and the smart album APP module is also interactively connected to a speaker and a display screen or a music platform server.
3. The intelligent photo album system supporting voice-on-demand music according to claim 2, characterized in that: The speaker and the display screen are both arranged in the smart album terminal. The speaker is used to play the detected music, and the display screen is used to display the dialogue text generated by the user voice conversion, as well as to play the retrieved music MV and display the lyrics.
4. The intelligent photo album system supporting voice-on-demand music according to claim 2, characterized in that: The microphone is used to collect voice information input by the user.
5. The intelligent photo album system supporting voice-on-demand music according to claim 1, characterized in that: The smart album server is also provided with a smart question-answering service module, and the smart question-answering service module is used to conduct interactive question-answering with the user.
6. The intelligent photo album system supporting voice-on-demand music according to claim 1, characterized in that: A large language model system module is provided in the large language model server, and the large language model system module is interactively connected with the intelligent question-answering service module.
7. The intelligent photo album system supporting voice-on-demand music according to claim 6, characterized in that: The large language model system module constructs a vector database, combines keyword search, semantic matching and graph reasoning, and uses a language-multimodal comparison model to map text and audio cross-modal data into a unified semantic space for accurate audio retrieval.
8. The intelligent photo album system supporting voice-on-demand music according to claim 1, characterized in that: A music platform system module is set in the music platform server, and the music platform system module is interactively connected with the large language model system module or the smart album APP module.
Citation Information
Patent Citations
AI multimedia circulating scrolling display intelligent touch household photo album
CN118409675A