Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

372results about "Metadata audio data retrieval" patented technology

Systems and methods for generating playlists by applying search prompts to a model configured to generate structured queries

An electronic device associated with a media-providing service stores, in a vector space, a plurality of respective vector representations for respective media content items. The electronic device receives a user input, including a text string. The electronic device generates, using a neural network, a structured query based on the text string. The electronic device determines, based on the structured query, whether to generate a vector representation of a portion of the text string. When the electronic device determines to generate the vector representation of the portion of the text string, it generates the vector representation of the portion of the text string, wherein the vector representation is embedded in the vector space, and identifies a set of media items using the vector representation of the portion of the text string. And the electronic device provides one or more select media items from the set of media items to a user.
Owner:SPOTIFY

Information retrieval method, related device, equipment and storage medium

The invention discloses an information retrieval method, a related device, equipment and a storage medium, and is applied to the field of artificial intelligence. The method comprises the steps of obtaining to-be-queried information and target instruction information; based on the to-be-queried information and the target instruction information, obtaining a first token sequence through a feature extraction layer included in an information retrieval model; performing bidirectional attention coding on the first token sequence through a backbone network included in the information retrieval model to obtain a second token sequence; pooling the second token sequence to obtain a target embedded vector; and according to the target embedding vector, obtaining a preset number of candidate information with the highest similarity as a retrieval result. According to the method, the complexity and the calculation cost of model deployment are reduced, and the target embedding vector generated by the model is more discriminative, so that the accuracy of a retrieval result is improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Audio and music score matching method and system

The invention relates to the technical field of music score matching, in particular to an audio and music score matching method and system, and the method comprises the steps: providing a reference database of a music score, which comprises multiple segments of expected playing characteristics of a plurality of music units; acquiring first note data generated by playing of the user at first time; judging whether second note data generated at adjacent second time before the first time is successfully matched with the first target music unit or not; if yes, determining a first dynamic prediction sequence at the first time according to the first target music unit; judging whether the corresponding music unit obtained in the first dynamic prediction sequence is a second target music unit of the first note data or not; and if yes, marking the second target music unit as the current playing node. According to the music score matching technology, the matching time axis can be dynamically adjusted in combination with the playing state, and then positioning matching of the audio and the music score can be rapidly achieved through the limited matching calculation amount.
Owner:GRANMUS STAFF TECHNOLOGIES (CHONGQING) CO LTD

Maintenance guidance method and device for household appliances

The invention provides a maintenance guidance method and device for a household electrical appliance. The method comprises the following steps: constructing a unified vector database in advance based on text data, image data and audio data related to a household electrical appliance fault; acquiring field condition information of the household electrical appliance, wherein the field condition information comprises video information, audio information and / or text information; respectively carrying out image matching, audio matching and / or text matching on the field condition information based on the unified vector database, and outputting knowledge fragments as retrieval results; and inputting a retrieval result into the large language model, and outputting a maintenance guidance scheme of the household electrical appliance. According to the method, a retrieval and reasoning framework based on multi-modal feature direct fusion can be constructed, information dimension reduction loss is avoided by directly analyzing image, audio and text features of faults, and the comprehensive accuracy of fault diagnosis is improved to 95% or above; and through a dynamic knowledge retrieval and enhanced generation technology, the risk of'knowledge illusion 'possibly occurring in the professional field of the general large language model is effectively reduced.
Owner:HANGZHOU ROBAM APPLIANCES CO LTD

Music education practice system and method thereof

The invention discloses a music education practice system and a method thereof, and relates to the technical field of education. The system comprises a music score library, a teacher unit, a student unit and a display unit. Wherein the music score library is used for storing online open source music scores; the teacher unit is used for carrying out music score analysis on the music score and splitting the music score into corresponding signal components according to a singing mode and a playing mode to form an original audio; the student unit is used for practicing target music through peripheral equipment and recording by taking the length of each piece of music as a cycle to obtain a practicing audio; and the display unit is used for dynamically displaying the original audio and the practice audio in the form of visual notes, and deducting scores when the original audio and the practice audio have signal inconformity to obtain a final practice scoring result. According to the invention, the user can be assisted to improve the convenience of music practice.
Owner:东营职业学院

Music generation method based on multi-modal cross attention and dynamic feature fusion

The invention discloses a music generation method based on multi-modal cross attention and dynamic feature fusion, which is characterized by comprising the following steps of: extracting semantic features of a text by using a pre-training language model, and extracting image region features through a pre-training visual model; a multi-head cross attention module is used, text features are used as Query, image features are used as Key and Value, and a guide module pays attention to features, strongly related to texts, in pictures; constructing a small neural network prediction modal weight, and dynamically weighting and fusing text features and picture features; and extracting emotion information in the text and the picture by using a long-short-term memory network and a convolutional neural network, and guiding a diffusion model to generate symbol music according to the fusion features. Compared with the prior art, the method has the advantages of efficiently capturing potential information of texts and pictures, being high in cross-modal information fusion capability, enhancing the consistency of potential emotions of music and input information and the like, can improve the matching degree of music and input to a certain extent, and has a good application prospect.
Owner:EAST CHINA NORMAL UNIV +1

Language model augmented audio selection and generation

The present disclosure relates to a system and method for selecting and generating audio using a large language model. The method includes receiving from a user a text-based prompt for a desired song, generating a song specification from a prompt that includes the text-based prompt and instructions on how to create a suitable instruction file format for representing the requested song, for each of the list of tracks in the song specification, generating a ranked list of potential sound loops matching the song specification for a selected track, selecting a sound loop from the ranked list of potential sound loops for each of the list of tracks, and generating a track specification file including the sound loop selected for each of the list of tracks.
Owner:OUTPUT INC

Media content management

A system and method for media content management include creating, via a digital vault, a container file comprising media content submitted by a first user and content metadata; verifying, via the digital vault, a completeness of the content metadata associated with the media content in the container file; classifying, via the digital vault, the container file based on the completeness of the media content; capturing, via the digital vault, event metadata when a second user gains access to the container file, the event metadata comprising at least one of identification of the second user, an activation timestamp, a duration of access, portions of the container file accessed, and changes to the container file; and enabling a private communication channel between parties affiliated with the media content to permit messaging among the parties affiliated with the media content via the private communication channel.
Owner:TUNEGO INC

Sound effect recommendation method, system and equipment based on user preference and scene perception

The invention relates to the technical field of Internet information processing, in particular to a sound effect recommendation method, system and device based on user preference and scene awareness, and the method comprises the steps: carrying out the deep analysis of collected user preference data, and constructing a user preference model; when an audio recommendation request of a user side is received, scene perception information matched with a current vehicle position is obtained, and a current scene is determined; matching the current scene with the user preference model to obtain a candidate sound effect set matched with the current scene and the user preference; checking each candidate sound effect data in the candidate sound effect set, screening the candidate sound effect data according to the sound effect data, and generating a recommendation result; and feeding back the recommended sound effect contained in the sound effect recommendation result. According to the scheme, the personalized requirements of the user can be fully considered, and the preferable sound effect recommendation information is provided for the user, so that the satisfaction degree of the user on sound effect recommendation is improved.
Owner:CHINA FAW CO LTD

Audio playing method and device with contexts kept continuous, equipment and storage medium

The invention relates to the technical field of audio playing, and discloses an audio playing method and device capable of keeping contexts continuous, equipment and a storage medium. Matching an audio context model conforming to a preset rule from a playing record database; analyzing the audio context model, and extracting audio playing state information and sound effect configuration information recorded in the audio context model; and adjusting the sound effect of the target playing device, and controlling audio playing according to the audio playing state information. Through the playing control mode, switching playing of multiple sound sources can be achieved, continuous playing after switching can also be achieved, and the problems that audio playing repetition is high after traditional switching, and the user experience feeling is reduced are solved.
Owner:LINKPLAY TECHNOLOGY INC NANJING

Multi-modal content based automated feature recognition

A system includes a computing platform having processing hardware, and a memory storing software code and a machine learning (ML) model-based feature classifier. When executed, the software code receives media content including a first media component corresponding to a first media mode and a second media component corresponding to a second media mode, encodes the first media component using a first encoder to generate multiple first embedding vectors, and encodes the second media component using a second encoder to generate multiple second embedding vectors. The software code further combines the first embedding vectors and the second embedding vectors to provide an input data structure for a neural network mixer, process, using the neural network mixer, the input data structure to provide feature data corresponding to a feature of the media content, and predict, using the ML model-based feature classifier and the feature data, a classification of the feature.
Owner:DISNEY ENTERPRISES INC

Media identification system

A media identification system is provided. The system comprises an audio input configured to receive an audio signal, and an audio clip extraction module configured to extract an audio clip from the audio signal. The system further comprises an audio clip processing module configured to generate metadata based upon the audio clip, and a first communication interface configured to transmit media identification data corresponding to the audio clip to a media identification server when the metadata based upon the audio clip meets a predetermined requirement, wherein the predetermined requirement comprises the metadata indicating that the audio clip comprises music.
Owner:AUDOO LTD

Information processing system, information processing method, and computer program

The present invention provides an information processing system for performing processing related to music using AI technology. An information processing system according to the present invention includes an acquisition unit that acquires a trained music foundation model, and a generating unit that generates a trained model adapted to a downstream task on the basis of a training dataset related to the music foundation model and the downstream task. The model generating unit generates a plurality of trained models for each downstream task, on the basis of common intermediate features with the music foundation model, and a plurality of training datasets respectively corresponding to the plurality of the downstream tasks.
Owner:SONY GROUP CORP

Method and apparatus for determining multimedia content, electronic device, and storage medium

Provided are a method and apparatus for determining multimedia content, an electronic device, and a storage medium. The method comprises: acquiring a target material and text information (S110); determining a material feature, a text feature, and an intent type on the basis of the target material and the text information by means of a preset model (S120); and according to the intent type, determining corresponding target multimedia content for the target material on the basis of at least one target object in a target object set (S130), wherein the target object set comprises keywords in the text feature, a text vector corresponding to the text feature, and a material vector corresponding to the material feature.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Systems and methods for transforming digital audio content

A system for platform-independent visualization of audio content, in particular audio tracks utilizing a central computer system in communication with user devices via a computer network. The central system utilizes various algorithms to identify spoken content from audio tracks and identifies “great moments” and / or selects visual assets associated with the identified content. Audio tracks, for example Podcasts, may be segmented into topical audio segments based upon themes or topics, with segments from disparate podcasts combined into a single listening experience, based upon certain criteria, e.g., topics, themes, keywords, and the like.
Owner:TREE GOAT MEDIA LLC

Interrogation record and audio and video recording linkage method and system

The invention provides an interrogation record and audio and video recording linkage method and system, and belongs to the technical field of information processing, and the method comprises the steps: obtaining interrogation record text data and segmented video recording data, and respectively labeling category labels; performing semantic extraction on the record text data to obtain semantic feature information, and converting the feature information into text semantic feature vectors; frame extraction is carried out on the video data, image features of each frame of the video data are extracted, similarity screening is carried out, and the image features are converted into key frame semantic feature vectors; constructing a cross-modal embedding space model, respectively encoding the text semantic feature vector and the key frame semantic feature vector of the category label through two encoders, then splicing and fusing to obtain a fused feature vector, converting and mapping the fused feature vector into a sharing layer, completing feature association of the record text data and the segmented video recording data, and obtaining the segmented video recording data. And subsequently, cross-modal search is carried out. According to the method, cross-modal information association is realized through association of the semantic feature vectors of the record and video recording information, and a basis is provided for subsequent rapid calling, research and judgment.
Owner:SHAANXI POLICE VOCATIONAL COLLEGE (SHAANXI POLITICAL & LEGAL MANAGEMENT CADRE COLLEGE)

Multi-language interactive learning system based on speech recognition

The invention relates to the field of voice signal processing, in particular to a multi-language interactive learning system based on voice recognition, which is characterized in that sound waves and mouth shape images are synchronously acquired and discretized by the system, and cross-modal coding is formed after alignment; secondly, the code is injected into a micro-ring photon reserve network through phase modulation to unfold time sequence characteristics, the code is mapped into a quaternion graph to be embedded, and segment boundaries are extracted through a diffusion-pulse coupling method; then, according to a fixed field sequence, encapsulating the quaternion graph embedding and segmentation data into an object mark prompt, inputting the object mark prompt into a low-rank adaptive language model, and generating semantic segmentation data and text transcription; and finally, the adaptive learning module implements sparse gradient updating on the low-rank weight and pulse network by using an integer fractal hash exclusive-or difference mask, and adopts support-query element learning for synchronous iteration after user clarification. According to the system, low-power-consumption, high-precision and second-level accent self-adaptive multi-language voice interaction is realized on the end side.
Owner:SICHUAN COLLEGE OF ARCHITECTURAL TECH

Enterprise platform with integrated user-curated playlist

A merchant processing system accesses a user interface of a content delivery service, to identify media content items. In response to merchant inputs, the merchant processing system causes generation of a playlist of media content items in a playlist format of the content delivery service. In response to determining that the user is accessing an application of the merchant, the merchant processing system causes display of the playlist to the user on a device associated with the user, within one or more screen displays of the application. The merchant processing system then receives a playback request representing a selection by the user of a media content item on the playlist, in response to the display of the playlist to the user, and causes transmission of the selected media content item to the device associated with the user, to cause playback of the selected media content item, based on the playback request.
Owner:BLOCK INC

Music education resource management method and system

The invention relates to the technical field of education management, in particular to a music education resource management method and system, and the method comprises the steps: carrying out the preprocessing of music resource data, constructing music resource classification, obtaining various types of music resource sub-data, building a management model, and inputting the music resource sub-data and user data resources into the management model. The method comprises the steps of screening representative users and representative music through a management model, establishing a coupling relationship between the representative users and the representative music, training the management model based on the coupling relationship between the representative users and the representative music, finally obtaining real-time user data, and adjusting playing training of the real-time users based on the real-time representative music output by the trained management model. According to the method, the representative users and the representative music are screened out, so that the coupling relation between the representative users and the representative music is established, the management model can perform music recommendation for the real-time users according to the real-time user identity data, and the accuracy of music recommendation for the users is improved.
Owner:HEBEI VOCATIONAL COLLEGE OF FOREIGN LANGUAGES

Speech recognition device and operating method thereof

Provided are a method and device for speech recognition. The speech recognition method includes: receiving a speech signal generated by an utterance of a user; identifying a named entity from the received speech signal; determining a speech signal portion, which corresponds to the identified named entity, from the received speech signal; generating a first acoustic embedding vector corresponding to the speech signal portion, based on an acoustic embedding model; determining a second acoustic embedding vector that is one of a plurality of acoustic embedding vectors corresponding to a plurality of named entities included in an acoustic embedding database (DB), based on distances between the plurality of acoustic embedding vectors and the first acoustic embedding vector; determining a corrected named entity corresponding to the second acoustic embedding vector; and providing a result of speech recognition with respect to the speech signal, based on the corrected named entity.
Owner:SAMSUNG ELECTRONICS CO LTD

Time code acquisition method, time code acquisition device, electronic equipment and program product

The invention is suitable for the technical field of multimedia, and provides a time code acquisition method, a time code acquisition device, electronic equipment and a program product. The method comprises the following steps: acquiring a plurality of materials; under the condition that a first material exists in the multiple materials and the first material contains audio data, if a second material exists in the multiple materials, determining a time code of the first material based on a time code of the second material; wherein the first material is a material which does not contain the time code or the time code is wrong, the second material is a material which contains the audio data and the time code, and the audio waveform of the contained audio data is matched with the audio waveform of the audio data of the first material. According to the method and the device, the time code of the material which does not contain the time code or has the wrong time code can be obtained, so that the availability of the material is improved.
Owner:APUTURE IMAGING IND CO LTD

Engine abnormal sound detection method, device, equipment and medium

The invention discloses an engine abnormal sound detection method, device and equipment and a medium. The engine abnormal sound detection method comprises the following steps: acquiring actual noise audio data of a to-be-detected engine; processing the actual noise audio data to generate actual noise audio feature data; establishing a standard audio database; wherein the standard audio database comprises standard noise audio feature data under various combination conditions of various influence parameters; determining matched noise audio feature data according to the feature information of the to-be-detected engine and the standard noise audio feature data; according to similarity parameters of the matched noise audio feature data and the actual noise audio feature data, abnormal sound faults of the engine are recognized; and when the abnormal sound fault of the engine is identified, the abnormal sound fault type of the engine is diagnosed according to the actual noise audio feature data.
Owner:FAW JIEFANG AUTOMOTIVE CO

Audio production model file creation method and electronic equipment

The invention relates to an audio production model file creation method and electronic equipment. The method comprises the steps of obtaining an original audio file; analyzing the original audio file through a renderer, and generating corresponding preset metadata elements; wherein the preset metadata elements comprise an audio program element, an audio content element, an audio object element, an audio track unique identification element, an audio packet format element, an audio channel format element, an audio stream format element and an audio track format element; based on the relation between the preset metadata elements, establishing a reference relation of the corresponding elements; generating a metadata file from the preset metadata elements according to the preset metadata elements and the reference relationship; creating a channel allocation file for connecting the metadata file and an audio track of the original audio file; and establishing a guide relationship between the metadata file and the channel allocation file, and generating an audio production model file.
Owner:SINE MICRO (BEIJING) ELECTRONIC TECH CO LTD

Methods and apparatus to identify media

Methods, apparatus, systems and articles of manufacture are disclosed to identify media. An example method includes: in response to a query, generating an adjusted sample media fingerprint by applying an adjustment to a sample media fingerprint; comparing the adjusted sample media fingerprint to a reference media fingerprint; and in response to the adjusted sample media fingerprint matching the reference media fingerprint, transmitting information associated with the reference media fingerprint and the adjustment.
Owner:GRACENOTE INC

Playlist Selection for Audio Streaming

An example embodiment may involve determining that a client device (such as a smartphone, tablet, or in-automobile audio device) is in an automobile and that the client device has access to a playlist of audio content. Possibly based on the client device being in the automobile and having access to the playlist of audio content, the client device may request a stream of the audio content. As a consequence of making the request, the client device may receive the stream of the audio content and begin audible playout of the audio content.
Owner:GRACENOTE DIGITAL VENTURES LLC

Electronic device stores tag information of content

An electronic device according to an embodiment comprises a memory, a display, and a processor operatively connected to the memory and the display, wherein the processor may be configured to: collect speech data; match the collected speech data with user information related to the collected speech data and store, in the memory, association information between the collected speech data and the user information; when generating content, detect speech data of the content input that is input during generation of the content; and when there is user information matching with the detected speech data in the memory, store the user information matching with the detected speech data of the content as tag information of the content.
Owner:SAMSUNG ELECTRONICS CO LTD

Song semantic processing method, computer equipment and computer storage medium

The embodiment of the invention discloses a song semantic processing method, computer equipment and a computer storage medium. The multi-modal representation of the song is subjected to discretization processing to generate sparse semantic representation, the multi-modal representation of the song is subjected to dimension reduction mapping processing to obtain dense semantic representation, the sparse semantic representation represents coarse-grained content features of the song, and the dense semantic representation represents fine-grained content features of the song; multi-modal attributes of songs can be fully described through coarse and fine semantic dimensions, and the ability of the model to understand the content of new and cold songs is greatly improved. And obtaining user behavior representation generated based on the historical behaviors of the user, wherein the user behavior representation can represent song listening preference characteristics of the user. The multi-modal semantic representation of the song and the user behavior representation are fused, the semantic content of the song and the listening preference of the user can be more comprehensively and accurately captured, and the problem that the new cold song is difficult to accurately put due to lack of historical data is effectively solved.
Owner:TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD

Techniques for audio track analysis to support audio personalization

Techniques for enabling personalization of audio tracks include selecting a portion of an audio track that is representative of the audio category, creating an audio sample from the portion of the audio track, playing the audio sample for a user, and adjusting, based on an input from the user while the audio sample is playing, a personalization setting for the user to be used when playing back audio from the audio category.
Owner:HARMAN INT IND INC

Blockchain-based clinic epidemic monitoring method, device, medium and equipment

The application discloses a clinic epidemic situation monitoring method and device based on a blockchain, a medium and equipment, and belongs to the technical field of blockchains. The method comprises the following steps: monitoring a doctor in a clinic to obtain a monitoring video of a patient during a treatment process; converting the monitoring video into monitoring audio; detecting whether the patient has symptoms of an epidemic disease according to the monitoring audio; if the patient has symptoms of the epidemic disease, generating epidemic situation monitoring content according to the monitoring video or the monitoring audio; storing clinic information of the clinic, identity information of the patient and the epidemic situation monitoring content in a blockchain connected to the clinic; and sending a notification message to an epidemic prevention center so that the epidemic prevention center monitors the patient according to the notification message. The application automatically monitors the epidemic situation of the clinic, saves manpower, avoids the problems of missing reports, hidden reports and tampering with monitoring videos, and avoids wasting disk space. Sending the notification message to the epidemic prevention center can improve the response speed to the epidemic situation.
Owner:HANGZHOU RIVTOWER TECH CO LTD

Emotion and behavior anomaly detection and early warning method and device, equipment and storage medium

The invention discloses an emotion and behavior anomaly detection and early warning method and device, equipment and a storage medium, and the method employs a wearable detection device which can detect the speech and behavior of a server, the speech, expression and posture of a serviced person, and the environment where the serviced person is located. And the emotion and the holding behavior of the server and the serviced person are comprehensively analyzed, so that early warning is carried out on the severe emotion change and the violent behavior. Through a camera and a microphone on the multi-mode cooperative emotion sensing module, emotion changes and voice changes of a server and a serviced person can be clearly recorded, and the emotions and actions of the server and the serviced person are analyzed through the multi-mode cooperative emotion sensing module. On one hand, emotion and behavior changes of a server and a serviced person are recorded to provide data support for subsequent better services, and on the other hand, events can be better restored through the recorded data after accidents occur. The invention relates to a character and environment detection technology.
Owner:SOUTH CHINA UNIV OF TECH +1