Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

30results about "Metadata audio data retrieval" patented technology

Music database retrieval system based on feature extraction

PendingCN122196226AMetadata audio data retrievalBiological models
The application discloses a music database retrieval system based on feature extraction, and particularly relates to the technical field of music information retrieval, comprising three modules: an adversarial feature enhancement module, which eliminates the sound quality variation information in the features through gradient reversal adversarial training; a multi-level time sequence fusion module, which adopts hierarchical dilated convolution and a gated attention mechanism to fuse multi-scale time sequence features; and a cross-version contrast learning module, which combines a dynamic difficult example mining strategy to optimize the feature space distribution. The application cooperatively solves the problems of the prior art, such as sensitivity to audio quality changes, insufficient time sequence modeling, and weak cross-version generalization capability, and significantly improves the accuracy, robustness and practicality of the music retrieval system in complex real scenes.
Owner:QUJING NORMAL UNIV

A music recommendation method and device, a vehicle and a medium

PendingCN122112300AMetadata audio data retrievalInference methodsPersonalizationIn vehicle
The application relates to the technical field of vehicle-mounted music, and discloses a music recommendation method and device, a vehicle and a medium, the method comprising the following steps: obtaining atmosphere information, wherein the atmosphere information is used for representing the atmosphere of a user in a vehicle; determining the favorite weights of the user for various music types according to the atmosphere information, so as to obtain a favorite weight set; and matching corresponding target music from a music library through the favorite weight set. The application improves the accuracy of personalized recommendation of vehicle-mounted music.
Owner:CHONGQING CHANGAN AUTOMOBILE CO LTD

system

PendingJP2026098800AInput/output for user-computer interactionMetadata audio data retrieval
We provide the system. [Solution] Means for obtaining user image information, A means for analyzing the aforementioned image information to infer the atmosphere and emotional state, means for generating musical information based on the aforementioned atmosphere and emotional state, Means for transmitting the generated music information to the user's terminal device, A system that includes this.
Owner:SOFTBANK GROUP CORP

Song searching method, device, searching system and computer readable storage medium

PendingCN122196227AMetadata audio data retrievalSpecial data processing applications
The application discloses a song search method and device, a search system and a computer readable storage medium, and belongs to the technical field of search. The embodiment receives a song search request, wherein the song search request comprises song search information; searches a first song from a first storage space corresponding to a first search engine based on the song search information through the first search engine, wherein the first storage space stores song information of all songs in a music platform; searches a second song from a second storage space corresponding to a second search engine based on the song search information through the second search engine, wherein the second storage space stores song information of songs meeting preset conditions in the music platform; and generates a search result corresponding to the song search request based on the first song and the second song, so that the stability of the search system and the accuracy of the search result can be improved.
Owner:HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD

Audio-visual video semantic parsing method and system suitable for cross-media information retrieval

PendingCN122240879AMetadata audio data retrievalMetadata video data retrieval
This application relates to a semantic parsing method and system for audiovisual videos suitable for cross-media information retrieval. The method includes: extracting video frame sequences and audio slice sequences of equal length from the audiovisual video to be processed; mapping the video frame sequences and audio slice sequences to a shared feature space; constructing an audiovisual feature distance matrix based on the feature differences and temporal deviations between the video frame sequences and audio slice sequences; transforming the audiovisual feature distance matrix into an undirected fully bipartite graph; solving the graph matching algorithm to generate a minimum distance matching sequence; determining the audiovisual feature matching score based on the degree of temporal inversion of the minimum distance matching sequence and the corresponding values ​​of each node in the minimum distance matching sequence in the spatial cost matrix; and determining the feature storage path of the video frame sequences and audio slice sequences based on the comparison results of the audiovisual feature matching scores with a preset judgment benchmark. This application can improve the accuracy of cross-media retrieval.
Owner:SHAANXI GUOBO ZHENGTONG INFORMATION TECH CO LTD +1

Method and apparatus for generating meeting minutes, electronic device and medium

ActiveCN116756367BMetadata audio data retrievalSemantic analysisTelecommunicationsProcessing
The present disclosure provides a method, device, electronic device, computer readable storage medium and computer program product for generating meeting minutes, relates to the field of natural language processing, and particularly relates to the field of speech recognition and keyword extraction. The implementation scheme is: obtaining first conference audio, wherein the first conference audio includes information indicating that a request execution person is requested to execute a to-do item; obtaining second conference audio according to the first conference audio, wherein the second conference audio is associated with the execution person, and the second conference audio includes information indicating whether the execution person commits to execute the to-do item; and in response to determining that the execution person commits to execute the to-do item according to the second conference audio, generating a target meeting minutes based on the first conference audio, wherein the target meeting minutes indicate that the to-do item is executed by the execution person.
Owner:APOLLO INTELLIGENT CONNECTIVITY (BEIJING) TECH CO LTD +1

A method and system for automatically generating sales text videos

ActiveCN121071181BHigh originalityIncrease randomnessMetadata audio data retrievalTelevision system detailsEngineeringAudio frequency
This invention relates to the field of video generation technology, specifically disclosing a method and system for self-generating sales text videos. The method includes receiving sales text uploaded by a user; performing semantic analysis on the sales text to segment it into sentences and words in each sentence; creating a sales video based on the words; recognizing the sales video using AI to obtain recognized text; comparing the recognized text with the sales text to calculate text similarity; and reading audio from a preset audio library based on the text similarity and inserting it into the sales video. This invention performs semantic analysis on the sales text, segments it into words, creates a sales video based on the words, recognizes the sales video, compares the recognition results with the sales text, and adjusts the audio insertion part according to the comparison results, thereby optimizing the originality and randomness of the sales video.
Owner:TAIDOU TECH GRP CO LTD

An audio text correlation evaluation method based on a multi-expert model, an audio retrieval system, a storage medium and a program product

PendingCN122173674AMetadata audio data retrievalPattern recognitionSemantic alignment
The application belongs to the technical field of audio processing, and a plurality of pre-trained audio text expert models are introduced to construct a multi-expert audio text semantic representation. Meanwhile, the consistency relationship between audio and text at the overall semantic level and the semantic inconsistency between audio and text that is perceptually significant to humans are modeled through a semantic alignment branch and a semantic mismatch branch respectively, and a multi-branch correlation score fusion mechanism is used to output an audio text correlation prediction result that is highly consistent with human subjective evaluation. This method can simultaneously consider semantic consistency and perceptual differences, effectively achieving more accurate and more human subjective evaluation standard-compliant audio text correlation evaluation. Meanwhile, based on the above method, the application constructs an audio retrieval system that evaluates and sorts the correlation between a text query and audio samples in an audio library to realize an audio retrieval function oriented to natural language description.
Owner:HARBIN ENG UNIV

Methods and systems for playing back indexed conversations based on the presence of other people

ActiveUS12645717B2Metadata audio data retrievalComputer security arrangementsDialog systemSystem monitor
Methods and systems are provided herein for playing back indexed conversations based on the presence of other people. When a user asks a query, the system monitors the area, determines the other users in the area, and searches its database for a conversation that addresses the query in consideration of the other users present in the area. The system filters the indexed conversations to find conversations that included all the users present and determines the best matching conversation based on the words of the query as well as the keywords from the conversation. Once the system has determined the best match conversation, the system plays back the conversation to the user.
Owner:ADEIA GUIDES INC

Intelligent audio generation system, method, device and medium supporting multi-entity interaction

PendingCN122196223AMetadata audio data retrievalCo-operative working arrangements
The application discloses a kind of intelligent audio generation system, method, equipment and medium supporting multi-entity interaction, it is related to audio equipment field.System includes: multiple NFC cards with unique identifier and being defined as story element attribute;Intelligent audio device is used to detect and read the unique identifier of multiple NFC cards in induction area by NFC card reading module, then combination generates combination request instruction and sends to cloud server;Cloud server is used to receive combination request instruction, according to multiple unique identifiers in story logic rule base matching determines corresponding story line logic, retrieves audio segment from audio segment library and generates ordered audio segment playing sequence and issues;Intelligent audio device receives audio segment playing sequence and plays audio content therein in order.The application can generate differentiated audio content with logical association, break the single solidified interaction limit, and enrich the audio content interaction experience of user.
Owner:SHENZHEN WELLDY TECH CO LTD

Audio material auditing and storing method and system

ActiveCN121765111BImprove review efficiencyImprove reliabilityMetadata audio data retrievalBiological modelsEngineeringAudio frequency
The application discloses an audio material auditing and warehousing method and system, relates to the technical field of audio auditing, and has the technical scheme as follows: obtaining a target audio material to be audited, and performing compliance screening on the target audio material; if the compliance of the target audio material cannot be determined, determining a propagation scene label after analyzing metadata information and audio waveform features of the target audio material; if the propagation scene label belongs to a cross-channel propagation type, extracting cross-channel audit records of audio of the same type from a historical cross-channel audit database, and calculating a feature correlation coefficient between the historical audio and the target audio material; separating a to-be-determined audio segment from the target audio material, generating an audit report after judging the multi-channel compliance of the to-be-determined audio segment according to the feature correlation coefficient and a real-time audit rule library of each channel, and formulating an adaptive strategy to complete warehousing determination; and the effect is to provide strong support for the standardized management and effective propagation of audio materials.
Owner:HANGZHOU XIAOSHAN HUA NUMBER OF DIGITAL TV CO LTD +1

Lyric transcription systems, devices, and methods

PendingUS20260161705A1Metadata audio data retrievalSpecial data processing applications
A system is configured to facilitate lyric acquisition for audio content. The system accesses audio content and generates an AI-based lyric transcription that includes lyric segments comprising words. A user interface presents the lyric transcription in editable form to enable user modification and validation of words and segment boundaries. After user validation, the system generates an AI-based temporally aligned lyric transcription by determining a timestamp for each validated lyric segment based on the audio content. The temporally aligned lyric transcription is presented in editable form to enable user modification of timestamps and lyric text. The system receives user confirmation of finalized lyric segments and corresponding finalized timestamps. The system may construct a lyric transcription package comprising the finalized temporally aligned lyric transcription, a language designation, a track title, and an artist name, and may submit the package to distribution platforms.
Owner:MOISES SYSTEMS INC

Vehicle abnormal noise warning methods, devices, vehicles and storage media

ActiveCN116691549BTroubleshootingMetadata audio data retrievalSustainable transportationVehicle drivingData bank
This application provides a method, device, vehicle, and medium for alerting abnormal vehicle noises. The method, applied in the field of vehicles, includes: acquiring sound information generated by the vehicle, vehicle operation information, and vehicle driving information; determining the vehicle's operating scenario based on the vehicle operation information and driving information; comparing the sound information with normal sound information stored in a first database based on the operating scenario to generate a detection result; and outputting corresponding alert information according to the detection result and a corresponding alert mode. This method can alert users to vehicle noises based on their needs, enabling users to promptly detect vehicle malfunctions and ensuring vehicle driving safety.
Owner:GREAT WALL MOTOR CO LTD

Music knowledge graph construction method, electronic device, storage medium and air conditioner

ActiveCN115658914BMetadata audio data retrievalSemantic analysis
The application relates to a music knowledge graph construction method, an electronic device, a storage medium and an air conditioner. The method comprises the following steps: obtaining target music attributes of target music according to a target scene corresponding to a target air conditioner type, wherein the target music is music of a to-be-associated air conditioner type, and the target air conditioner type is an air conditioner type to be determined whether to match the target music; determining a target matching result according to the target music attributes, wherein the target matching result is used for indicating whether the target air conditioner type matches the target music; and associating the target air conditioner type with the target music according to the target matching result to obtain a music knowledge graph. The application solves the technical problem that directly using general music resources for music recommendation cannot recommend music suitable for different use scenes of air conditioners for users.
Owner:GREE ELECTRIC APPLIANCE INC OF ZHUHAI +1

Method for determining data representative of emergency situations, corresponding system and program.

PendingFR3169591A1Metadata audio data retrievalData processing applications
The invention relates to a method and system for determining emergency situations from audio signals received from a set of transmission channels used by a set of intervention teams, the method and system comprising at least the implementation of the steps: acquisition (S01) of an audio signal from a current transmission channel among the set of transmission channels; processing (S02) of the audio signal comprising: sampling at a predetermined frequency; segmentation of the sampled signal according to a predetermined interval, delivering a sequence of audio segments; and for each audio segment, extraction of spectral features, delivering a sequence of segments of spectral features; determination (S03) of a first data representative of an emergency situation of the current transmission channel as a function of the sequence of segments of spectral features of the audio signal.Fig 2.
Owner:STREAMWIDE

SONIFICATION OF NAVIGATION SEARCH RESULTS

UndeterminedDE102025115534B3Instruments for road network navigationMetadata audio data retrievalSound sourcesSonification
A method for sonifying search results, performed while a user is in the passenger compartment of a vehicle, comprises: receiving a response to a user-submitted query, the response containing a list of potential matches to the query, each potential match containing corresponding embedded metadata. For each potential match, the method further comprises: extracting the embedded metadata corresponding to the potential match; determining, based on the embedded metadata, a spatially located position within a playback sound field that the user can perceive as the sound source of the potential match; and outputting audio signals characterizing the potential match through a loudspeaker array in the passenger compartment to generate the playback sound field.The user perceives the potential hit as originating from the sound source at the spatially arranged point within the playback sound field.
Owner:GM GLOBAL TECHNOLOGY OPERATIONS LLC

Methods and apparatus to identify media based on historical data

PendingUS20260154336A1Metadata audio data retrievalSpeech analysisAlgorithmEngineering
Methods, apparatus, systems and articles of manufacture are disclosed to identify media based on historical data. An example method includes: comparing (a) a pitch shifted fingerprint, (b) a time shifted fingerprint, or (c) a resampled fingerprint to a reference fingerprint; in response to a match between any of (a) the pitch shifted fingerprint, (b) the time shifted fingerprint, or (c) the resampled fingerprint and the reference fingerprint, generating indications of (a) a pitch shift value, (b) a time shift value, or (c) a resample ratio that caused the match; in response to collecting broadcast media for a threshold period of time, processing the one or more indications; and in response to a request for a recommendation for information associated with a query, transmitting the recommendation including one or more frequencies of occurrence of (a) the pitch shift value, (b) the time shift value, or (c) the resample ratio.
Owner:GRACENOTE INC

Abnormal comment processing method, computer device and storage medium

ActiveCN116484045BMetadata audio data retrievalBiological models
The application relates to an abnormal comment processing method, computer equipment and a storage medium. The method comprises the following steps: performing clustering processing on a to-be-detected comment of a song to obtain a comment set corresponding to the song; confirming comment aggregation information of the comment set corresponding to the song, confirming comment aggregation information of the song according to the comment aggregation information of the comment set corresponding to the song; the comment aggregation information of the comment set is used for representing the overall aggregation degree of the comments in the comment set; sorting the songs according to the comment aggregation information of the songs and the song types of the songs; and performing corresponding abnormal comment processing on the sorted songs. The method can improve the efficiency of abnormal comment processing.
Owner:TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD

Vehicle-mounted music adaptive pushing method, vehicle and storage medium

PendingCN122196224AMetadata audio data retrievalTransmission
The application provides a vehicle-mounted music adaptive pushing method, a vehicle and a storage medium, and relates to the technical field of intelligent vehicles.The method comprises the following steps: firstly, based on the emotional state of a user in a vehicle and scene environment data, target content matched with the emotional state and the scene environment data is determined from a pre-constructed user music preference library.The scene environment data is used to represent the corresponding natural environment, time dimension and space travel scene of the user during the driving process of the vehicle.Secondly, based on the music playing state in the vehicle, the change characteristics of the emotional state and the interaction data of the user and the in-vehicle playing device, a pushing strategy of the target content is determined.The change characteristics are used to indicate the dynamic change attribute, fluctuation level and state switching characteristics of the emotional state of the user over time.Finally, based on the pushing strategy, the target content is pushed to the user.The technical problem of poor adaptability of vehicle-mounted music pushing in the related art is solved.
Owner:GREAT WALL MOTOR CO LTD

Media composition using non-fungible token (NFT) configurable pieces

ActiveUS12639407B2Metadata audio data retrievalFinanceMediaFLOEngineering
A system and method for receiving one or more non-fungible tokens (NFTs) that are associated with links to digital assets, ownership information, NFT metadata, media content metadata, or other media content information. The system may associate the one or more NFTs with other NFTs to create a collection of NFTs that form a song, album, video, or other collection / combination of media content. The system may provide a NFT collectible player to interact with the combination of NFTs in a particular order or for a particular duration.
Owner:TUNEGO INC

Enhanced ai-based audio-visual processing

PendingUS20260178642A1Metadata audio data retrievalSpeech recognitionVideo processingData source
Embodiments of the present disclosure relate to enhanced AI-based audio-visual processing. Various aspects integrate multimodal analysis functionality that seamlessly combines and / or selects from audio, video, and / or text data to provide a holistic understanding of multimedia content. Relative to existing technologies, such an approach enables more accurate and contextually aware interpretations by leveraging the full spectrum of available information in multimedia content. By integrating these disparate sources of data, various embodiments achieve a more nuanced analysis that captures the complexity and richness of real-world multimedia scenarios.
Owner:NVIDIA CORP

Music recommendation method, computer device and storage medium

The application relates to a music recommendation method, a computer device and a storage medium. The method comprises the following steps: acquiring a user portrait feature of a user to be recommended, and acquiring a music identifier sequence corresponding to the user to be recommended; the music identifier sequence is composed of music identifiers corresponding to the user to be recommended, music playing time latest and complete music playing music identifiers; the user portrait feature and the music identifier sequence are input into a pre-trained music recommendation model, and a user music preference feature of the user to be recommended is obtained through the music recommendation model; a target user is acquired from candidate users according to the user music preference feature of the user to be recommended; the user music preference feature of the target user is similar to the user music preference feature of the user to be recommended; a preferred music associated with the target user is acquired, and target music is acquired from the preferred music, and the target music is recommended to the user to be recommended. The method can improve the accuracy of music recommendation.
Owner:TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD

Systems and methods for transforming digital audio content

PendingUS20260155121A1Metadata audio data retrievalElectrophonic musical instrumentsUser deviceHuman–computer interaction
A system for platform-independent visualization of audio content, in particular audio tracks utilizing a central computer system in communication with user devices via a computer network. The central system utilizes various algorithms to identify spoken content from audio tracks and identifies “great moments” and / or selects visual assets associated with the identified content. Audio tracks, for example Podcasts, may be segmented into topical audio segments based upon themes or topics, with segments from disparate podcasts combined into a single listening experience, based upon certain criteria, e.g., topics, themes, keywords, and the like.
Owner:TREE GOAT MEDIA LLC

Data processing method, apparatus, medium, and computing device

ActiveCN117150072BMetadata audio data retrievalEnergy efficient computingAudio frequencyInformation retrieval
Embodiments of the present disclosure provide a data processing method, device, medium and electronic equipment, relating to the technical field of audio. The data processing method comprises: obtaining audio features of each first audio, and determining an audio group according to each audio feature; constructing second audio information according to first audio information of each first audio in the audio group, the second audio information being used to indicate a union set of each first audio information; and updating the first audio information of each first audio in the audio group to the second audio information. In the present disclosure, the second audio information representing the union set of each first audio information is obtained through the first audio information of the same audio, and the first audio information of each audio is then updated to the second audio information, that is, the information of each audio is completed with respect to the audio information to which the same audio belongs, without the need for manual searching for missing information of the audio, thereby improving the completion efficiency of the missing information of the audio.
Owner:HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD

Content retrieval method and apparatus, electronic device, and storage medium

PendingCN122220593AMetadata audio data retrievalNatural language data processing
The application provides a content retrieval method, device and equipment and a storage medium, which can be applied to the media field. The method comprises the following steps: acquiring a retrieval command input by an object; extracting first text data according to the retrieval command, wherein the first text data comprises information of retrieval content; splitting the first text data to obtain at least one key element according to a preset element type, wherein the preset element type is obtained according to a naming rule of a content type to which the retrieval content belongs; and matching the at least one key element with the name of at least one content in a content library to obtain target content meeting the retrieval command. The embodiment of the application is helpful to improve the accuracy of content retrieval.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Dynamic audio file generation

ActiveUS12639369B2Metadata audio data retrievalCommercePathPingData file
A method for time-sensitive data files in continuously generated data streams by a server computer comprises receiving user profile information associated with a plurality of different users and generating a dynamic data file that includes a plurality of user-selected metadata attributes, including a source link address to link with content via a universal resource locator file path. The method continues with the server computer receiving a request for a continuous data stream, generating the continuous data stream, inserting the dynamic data file in the continuous data stream and transmitting different content to a computing device associated with each of the plurality of different users.
Owner:IHEARTMEDIA MANAGEMENT SERVICES INC

Audio play method and system, and electronic device

ActiveUS12658180B2Metadata audio data retrievalSound input/output
An example method is discussed, which includes determining first audio data. The example method further includes determining a text file and metadata of the first audio data by performing content recognition processing on the first audio data. The example method further includes determining an audio type of the first audio data by performing audio type recognition processing on the first audio data based on the text file and the metadata. The example method further includes determining a sound effect mode corresponding to the audio type, where the sound effect mode is determined based on a preset correspondence between the audio type and the sound effect mode. The method further includes playing the first audio data based on the sound effect mode.
Owner:HUAWEI TECH CO LTD