Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

234results about "Metadata audio data retrieval" patented technology

Language model augmented audio selection and generation

The present disclosure relates to a system and method for selecting and generating audio using a large language model. The method includes receiving from a user a text-based prompt for a desired song, generating a song specification from a prompt that includes the text-based prompt and instructions on how to create a suitable instruction file format for representing the requested song, for each of the list of tracks in the song specification, generating a ranked list of potential sound loops matching the song specification for a selected track, selecting a sound loop from the ranked list of potential sound loops for each of the list of tracks, and generating a track specification file including the sound loop selected for each of the list of tracks.
Owner:OUTPUT INC

Multi-modal content based automated feature recognition

A system includes a computing platform having processing hardware, and a memory storing software code and a machine learning (ML) model-based feature classifier. When executed, the software code receives media content including a first media component corresponding to a first media mode and a second media component corresponding to a second media mode, encodes the first media component using a first encoder to generate multiple first embedding vectors, and encodes the second media component using a second encoder to generate multiple second embedding vectors. The software code further combines the first embedding vectors and the second embedding vectors to provide an input data structure for a neural network mixer, process, using the neural network mixer, the input data structure to provide feature data corresponding to a feature of the media content, and predict, using the ML model-based feature classifier and the feature data, a classification of the feature.
Owner:DISNEY ENTERPRISES INC

Media identification system

A media identification system is provided. The system comprises an audio input configured to receive an audio signal, and an audio clip extraction module configured to extract an audio clip from the audio signal. The system further comprises an audio clip processing module configured to generate metadata based upon the audio clip, and a first communication interface configured to transmit media identification data corresponding to the audio clip to a media identification server when the metadata based upon the audio clip meets a predetermined requirement, wherein the predetermined requirement comprises the metadata indicating that the audio clip comprises music.
Owner:AUDOO LTD

Enterprise platform with integrated user-curated playlist

A merchant processing system accesses a user interface of a content delivery service, to identify media content items. In response to merchant inputs, the merchant processing system causes generation of a playlist of media content items in a playlist format of the content delivery service. In response to determining that the user is accessing an application of the merchant, the merchant processing system causes display of the playlist to the user on a device associated with the user, within one or more screen displays of the application. The merchant processing system then receives a playback request representing a selection by the user of a media content item on the playlist, in response to the display of the playlist to the user, and causes transmission of the selected media content item to the device associated with the user, to cause playback of the selected media content item, based on the playback request.
Owner:BLOCK INC

Electronic device stores tag information of content

An electronic device according to an embodiment comprises a memory, a display, and a processor operatively connected to the memory and the display, wherein the processor may be configured to: collect speech data; match the collected speech data with user information related to the collected speech data and store, in the memory, association information between the collected speech data and the user information; when generating content, detect speech data of the content input that is input during generation of the content; and when there is user information matching with the detected speech data in the memory, store the user information matching with the detected speech data of the content as tag information of the content.
Owner:SAMSUNG ELECTRONICS CO LTD

Song semantic processing method, computer equipment and computer storage medium

The embodiment of the invention discloses a song semantic processing method, computer equipment and a computer storage medium. The multi-modal representation of the song is subjected to discretization processing to generate sparse semantic representation, the multi-modal representation of the song is subjected to dimension reduction mapping processing to obtain dense semantic representation, the sparse semantic representation represents coarse-grained content features of the song, and the dense semantic representation represents fine-grained content features of the song; multi-modal attributes of songs can be fully described through coarse and fine semantic dimensions, and the ability of the model to understand the content of new and cold songs is greatly improved. And obtaining user behavior representation generated based on the historical behaviors of the user, wherein the user behavior representation can represent song listening preference characteristics of the user. The multi-modal semantic representation of the song and the user behavior representation are fused, the semantic content of the song and the listening preference of the user can be more comprehensively and accurately captured, and the problem that the new cold song is difficult to accurately put due to lack of historical data is effectively solved.
Owner:TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD

Blockchain-based clinic epidemic monitoring method, device, medium and equipment

The application discloses a clinic epidemic situation monitoring method and device based on a blockchain, a medium and equipment, and belongs to the technical field of blockchains. The method comprises the following steps: monitoring a doctor in a clinic to obtain a monitoring video of a patient during a treatment process; converting the monitoring video into monitoring audio; detecting whether the patient has symptoms of an epidemic disease according to the monitoring audio; if the patient has symptoms of the epidemic disease, generating epidemic situation monitoring content according to the monitoring video or the monitoring audio; storing clinic information of the clinic, identity information of the patient and the epidemic situation monitoring content in a blockchain connected to the clinic; and sending a notification message to an epidemic prevention center so that the epidemic prevention center monitors the patient according to the notification message. The application automatically monitors the epidemic situation of the clinic, saves manpower, avoids the problems of missing reports, hidden reports and tampering with monitoring videos, and avoids wasting disk space. Sending the notification message to the epidemic prevention center can improve the response speed to the epidemic situation.
Owner:HANGZHOU RIVTOWER TECH CO LTD

Social Media Queue Across Multiple Streaming Services

Embodiments described herein may involve a social queue for use by a group of two or more media playback systems. An example method involves receiving, from a first media playback system, a first message indicating a first set of media items and receiving, from a second media playback system, a second message indicating a second set of media items. The method also involves generating a playback queue (i.e., a social queue) that includes the first set of media items indicated in the first message and the second set of media items indicated in the second message. The method may then involve transmitting, to at least one of the first media playback system and the second media playback system, the generated playback queue.
Owner:SONOS INC

Music database retrieval system based on feature extraction

The application discloses a music database retrieval system based on feature extraction, and particularly relates to the technical field of music information retrieval, comprising three modules: an adversarial feature enhancement module, which eliminates the sound quality variation information in the features through gradient reversal adversarial training; a multi-level time sequence fusion module, which adopts hierarchical dilated convolution and a gated attention mechanism to fuse multi-scale time sequence features; and a cross-version contrast learning module, which combines a dynamic difficult example mining strategy to optimize the feature space distribution. The application cooperatively solves the problems of the prior art, such as sensitivity to audio quality changes, insufficient time sequence modeling, and weak cross-version generalization capability, and significantly improves the accuracy, robustness and practicality of the music retrieval system in complex real scenes.
Owner:QUJING NORMAL UNIV

Inspection report generation system, method and device, computer equipment and storage medium

The invention relates to an inspection report generation system, method and device, computer equipment and a storage medium. The system comprises a server, at least one terminal and recording devices corresponding to the terminals, the recording devices are used for recording dialogues between doctors and current patients in real time and sending recorded audio data streams to the corresponding terminals, and the server is used for establishing corresponding data transmission channels for the terminals and transmitting the data transmission channels to the terminals. The terminal is used for uploading the received audio data stream to the server in real time through the data transmission channel, and the server is used for performing character recognition on the audio data stream to obtain text information and generating an examination report corresponding to a current patient according to the text information. By adopting the method, the audio data can be ensured not to be disordered in the process from acquisition to transmission, a reliable data basis is provided for subsequent transcription and report generation, and the accuracy of a check report is improved.
Owner:BEIJING UNITED FAMILY HOSPITAL CO LTD

System and method for discovering hit songs in a foreign language and popularizing those songs in listeners' native language music markets

The system's methodology combines a variety of information and communication technologies (ICT) tied to global immigration patterns—specifically connecting native and expatriate populations—to create a cross-cultural, “fusion music” listening experience. The system will foster this experience by creating music-based social networks that function as feedback loops between native and expatriate communities, generating crowd-sourced ‘music intelligence’—a means to identify hit songs in listeners' native languages and the promote, and accelerate, the popularity of these songs overseas in myriad foreign-language music markets. To promote songs worldwide, the system will track streams and rank songs (and podcasts) by the number of times listeners stream them, concurrently and asynchronously, on native and foreign-language platforms, and in a multiplicity of languages.
Owner:KAZANG INC

Information processing method and apparatus

This application discloses an information processing method and apparatus, applied to a first electronic device. The method includes: acquiring object information of a target object and content information indicated by the target object; generating a target information code matching the content information based on the object information of the target object and the content information indicated by the object information; splitting the target information code into N sub-information codes and displaying the N sub-information codes, so that a second electronic device can identify the N sub-information codes to obtain the content information; wherein, N is an integer greater than 1, the N sub-information codes include a first sub-information code, the first sub-information code is associated with the object information of the target object, and different sub-information codes are associated with different information.
Owner:VIVO MOBILE COMM CO LTD

Systems and methods for determining descriptors for media content items

An electronic device obtains a plurality of collections of media content items, each collection of media content items being associated with text. Based on how frequently a first media content item co-occurs with a first descriptor in text for respective collections of media items that include the first media content item, the electronic device generates, without user input, a new collection of media content items for a first user. The new collection of media content items corresponds to the first descriptor and includes the first media content item. The electronic device presents the new collection of media content items to the first user as a recommendation.
Owner:SPOTIFY

Information processing apparatus, information processing method, and program

To reduce labor for finding sound data related to a character string designated by a user.SOLUTION: An information processing device 1 includes a receiver 131 that receives a search string designated by a user, a searcher 135 that specifies a plurality of pieces of audio data in which a degree of association with the search string satisfies a predetermined condition, the degree of association indicating a relationship between a string and a sound, a map generator 136 that generates a map corresponding to a degree of similarity indicating a similarity between sounds of the plurality of pieces of audio data, and an outputter 137 that causes an information terminal to display a second area including a map corresponding to the degree of similarity together with a first area including a list of the plurality of pieces of audio data corresponding to the degree of association.SELECTED DRAWING: Figure 2
Owner:KDDI AGILE DEV CENT CO LTD

Song recommendation model training method, computer device, and storage medium

The application relates to a song recommendation model training method, computer equipment and a storage medium. The method comprises the following steps: taking song audio information and first text description information of a first sample song and second text description information of a second sample song as training samples, training a song feature extraction model in a contrast learning mode, obtaining text description features and song audio features of a third sample song through the trained song feature extraction model, splicing the text description features and the song audio features through a song recommendation model to obtain song fusion features of the third sample song, and training a song recommendation model by using the song fusion features of the third sample song and corresponding positive sample songs and negative sample songs of the third sample song to obtain a trained song recommendation model. The method can enable the song recommendation model to simultaneously analyze whether a song needs to be recommended to a user from the angles of audio and text description, and improve the accuracy of song recommendation.
Owner:TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD

3D retrieval device and 3D retrieval method

To provide a technique capable of improving retrieval accuracy of 3D information desired by a user.SOLUTION: A 3D retrieval device 1 that retrieves 3D includes an arithmetic unit 3 and a storage unit 4, and the arithmetic unit 3 extracts a 3D feature vector and a metadata feature vector from 3D and metadata added to the 3D, respectively, generates a fusion vector by fusing the 3D feature vector and the metadata feature vector, stores the fusion vector in the storage unit 4 as a dataset, and extracts a fusion vector having a high degree of association with a fusion vector of 3D serving as a retrieval query from the dataset.SELECTED DRAWING: Figure 1
Owner:HITACHI LTD

Query information rewriting model training method, music searching method, equipment and medium

The invention relates to a query information rewriting model training method, a music searching method, equipment and a medium. The training method comprises the following steps: acquiring sample music query information of a sample user and user preference of the sample user for a sample music search result; the sample music query information carries emoticon information, and the sample music search result is a search result obtained by searching through the sample music query information; inputting the sample music query information and the user preference into a query information rewriting model to be trained, and rewriting emoticon information carried in the sample music query information through the query information rewriting model to obtain prediction query rewriting information corresponding to the sample music query information; and training the query information rewriting model by using the difference between the predicted query rewriting information and the actual query rewriting information to obtain a trained query information rewriting model. By adopting the method, the application range of the query information rewriting model can be expanded.
Owner:YEELION ONLINE NETWORK TECH BEIJING

Data processing method and device, equipment, storage medium and computer program product

The embodiment of the application provides a data processing method, device, equipment, storage medium and computer program product, which can be applied to various fields or scenes such as artificial intelligence, block chain, cloud technology, intelligent transportation, smart home, vehicle-mounted and the like, wherein the method comprises: obtaining audio to be detected of an audio device to be detected; processing the audio to be detected to determine an audio feature set of the audio to be detected, the audio feature set comprising one or a combination of both of a first modal feature set and a second modal feature set; the first modal feature set is obtained by processing a frequency spectrum diagram of the audio to be detected, and the second modal feature set is obtained by processing at least two audio segments obtained by segmenting the audio to be detected; and determining a state detection result of the audio device to be detected according to the audio feature set. Through the embodiment of the application, the state detection result of the audio device can be quickly and accurately determined, and the method is simple and efficient.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Video processing method, mobile terminal and readable storage medium

This application discloses a video processing method, a mobile terminal, and a computer-readable storage medium. The method includes acquiring an original video to be processed and acquiring image content from the original video; filtering out matching background music based on the image content; and synthesizing the background music with the original video to output a target video with background music. This application makes the process of matching music to videos more intelligent, with almost all operations being completed automatically by the program. Furthermore, the background music matched in this application is based on the image content of the video, making the matched background music more accurate.
Owner:SHANGHAI TRANSSION CO LTD

A music recommendation method and device, a vehicle and a medium

The application relates to the technical field of vehicle-mounted music, and discloses a music recommendation method and device, a vehicle and a medium, the method comprising the following steps: obtaining atmosphere information, wherein the atmosphere information is used for representing the atmosphere of a user in a vehicle; determining the favorite weights of the user for various music types according to the atmosphere information, so as to obtain a favorite weight set; and matching corresponding target music from a music library through the favorite weight set. The application improves the accuracy of personalized recommendation of vehicle-mounted music.
Owner:CHONGQING CHANGAN AUTOMOBILE CO LTD

Music generator

Techniques are disclosed relating to determining composition rules, based on existing music content, to automatically generate new music content. In some embodiments, a computer system accesses a set of music content and generates a set of composition rules based on analyzing combinations of multiple loops in the set of music content. In some embodiments, the system generates new music content by selecting loops from a set of loops and combining selected ones of the loops such that multiple ones of the loops overlap in time. In some embodiments, the selecting and combining loops is performed based on the set of composition rules and attributes of loops in the set of loops.
Owner:AIMI INC

A large model-based voice dialogue retrieval method, device and medium

The application discloses a large model-based voice dialogue retrieval method, device and medium, and belongs to the technical field of voice retrieval. The method comprises the following steps: constructing a to-be-retrieved database and an inverted index database, wherein the former stores a text field, and the latter converts a phonetic alphabet. A user is guided to input a retrieval voice by using a preset dialogue template. The voice is converted into a phonetic alphabet retrieval information by using a voice recognition large model, and then a fuzzy matching algorithm is used to compare the inverted index and filter high matching degree fields. If the matching fields are many, a filtering voice is generated to further inquire; if the matching fields are few or none, the user is prompted to re-input or correct the input; and if the matching degree is 1, the retrieval field is directly determined. According to the field, the result in the to-be-retrieved database is found, and a unique direct reply is given; if the result is not unique, a missing field is identified and inquired, so that accurate retrieval is realized. The whole process is guided by a dialogue template. The application realizes the effects of flexible retrieval mode and low cold start cost to a certain extent by the above method.
Owner:SHANDONG SYNTHESIS ELECTRONICS TECH

system

We provide the system. [Solution] Means for obtaining user image information, A means for analyzing the aforementioned image information to infer the atmosphere and emotional state, means for generating musical information based on the aforementioned atmosphere and emotional state, Means for transmitting the generated music information to the user's terminal device, A system that includes this.
Owner:SOFTBANK GROUP CORP

AI-based digital collection background music recommendation method

PendingCN121479009AMetadata audio data retrievalBiological modelsCosine similarityRegularization algorithm
The invention discloses an AI-based digital collection background music recommendation method, and relates to the technical field of AI. The method comprises the following steps: acquiring a key frame average image, a dominant tone vector and a style label of a digital collection, and candidate music samples and audio signals of a music library; extracting dynamic features of the digital collections to generate rhythm representation vectors; processing the music audio to generate a rhythm structure vector, and calculating the rhythm similarity between the rhythm structure vector and the rhythm structure vector through a dynamic time warping algorithm; generating a digital collection visual style vector and a music style vector, and calculating a style consistency score by using a cosine similarity algorithm; an auxiliary music feature vector is generated through the digital collection visual style vector, and a guide score is calculated in combination with the bottom layer audio features; and calculating a comprehensive score index according to the rhythm similarity, the style consistency score and the guide score. According to the invention, background music recommendation is carried out on the digital collections through the comprehensive scoring index.
Owner:湖北云雷信息技术有限公司

Song searching method, device, searching system and computer readable storage medium

The application discloses a song search method and device, a search system and a computer readable storage medium, and belongs to the technical field of search. The embodiment receives a song search request, wherein the song search request comprises song search information; searches a first song from a first storage space corresponding to a first search engine based on the song search information through the first search engine, wherein the first storage space stores song information of all songs in a music platform; searches a second song from a second storage space corresponding to a second search engine based on the song search information through the second search engine, wherein the second storage space stores song information of songs meeting preset conditions in the music platform; and generates a search result corresponding to the song search request based on the first song and the second song, so that the stability of the search system and the accuracy of the search result can be improved.
Owner:HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD

Systems and methods for leveraging metadata for cross product playlist addition via voice control

Systems and methods for generating a playlist of audio content for a vehicle are disclosed. An audio input is received at a vehicle entertainment system of the vehicle while an audio content item is currently played by the vehicle entertainment system. The audio input includes an audio command trigger and an audio playlist command. In response to detecting the audio command trigger in the audio input, the audio input is parsed to determine the audio playlist command. A metadata associated with the audio content item is determined. In response to determining the audio playlist command, the audio content item is caused to be added to an audio content playlist of a third-party service based on the metadata of the audio content item.
Owner:ADEIA GUIDES INC

System and method for in-vehicle voice calls

To provide systems and methods for in-vehicle voice calls.SOLUTION: Embodiments are disclosed for providing voice calls to users of a motor vehicle. As an example, a method comprises, in response to a voice call, routing the voice call to at least one phone zone of a plurality of phone zones based on at least one of a user input and a source of the voice call, the plurality of phone zones being included in a cabin of a motor vehicle. In the method, sonic interference with a voice call may be reduced, and a main system audio may continue to play for unselected phone zones.SELECTED DRAWING: None
Owner:HARMAN INT IND INC