Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

317results about "Audio data clustering/classification" patented technology

Music segment tagging, sharing, and image generation

A method of automated generation of contextually-relevant images for a music segment includes receiving at least one of basic metadata information and lyric information for the music segment, generating a first prompt for a computer-implemented machine-learning language model based on the at least one of the basic metadata information and the lyric information, receiving context information from the computer-implemented machine-learning language model in response to the first prompt, generating a second prompt for the computer-implemented machine-learning language model based on the context information, generating a third prompt by providing the second prompt as an input to the computer-implemented machine-learning language model, and generating an image descriptive of the music segment by providing the third prompt as an input to a computer-implemented machine-learning image generation model.
Owner:HOOK MEDIA LLC

Methods and systems for speech emotion retrieval via natural language prompts

Methods and systems for generating training data for training a contrastive language-audio machine-learning model. A plurality of audio segments are retrieved from a speech emotion recognition (SER) database along with metadata associated with the audio segments. The metadata of each audio segment includes an emotion class. Words or terms associated with emotions are retrieved from a lexicon. A large language model (LLM) is executed on (i) the classes of emotion associated with the audio segments and (ii) the words or terms from the lexicon. This generates a plurality of text captions associated with emotion, which are stored in a caption pool. For each audio segment retrieved from the SER database, that audio segment is paired with one or more of the text captions from the caption pool that were generated based on the emotion class associated with that audio segment. This yields audio-text pairs for training a contrastive learning model.
Owner:ROBERT BOSCH GMBH

Method and apparatus for generating motion of virtual character, and method and apparatus for constructing motion library of virtual character

PendingUS20250278881A1Semantic analysisSpeech analysis
A method and an apparatus for generating a motion of a virtual character, and a method and an apparatus for constructing a motion library of a virtual character are provided, and belong to the field of computer technologies. The method for generating a motion of a virtual character includes: obtaining audio and text of a virtual character, the text indicating semantic information of the audio (201); determining a semantic tag of the text based on the text (202); retrieving a motion category matching the semantic tag and motion data belonging to the motion category from a preset motion library (203); and generating a motion sequence of the virtual character based on the motion data (204).
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Adaptive sample selection for data item processing

Methods, systems, and apparatuses, including computer programs encoded on computer storage media, for receiving a query relating to a data item that includes multiple data item samples and processing the query and the data item to generate a response to the query. In particular, the described techniques include adaptively selecting a subset of the data item samples using a selection neural network conditioned on features of the data item samples and the query. Then processing the subset and query using a downstream task neural network to generate a response to the query. By adaptively selecting the subset of data item samples according to the query, the described techniques generate responses to queries that are more accurate and require less computation resources than would be the case using other techniques.
Owner:GOOGLE LLC

Method and apparatus for classifying generated speech

A method and apparatus for classifying generated speech are disclosed. The method for classifying generated speech includes: applying a one-dimensional convolution operation to raw speech data to embed the raw speech data into a feature space and extract a feature vector; quantizing the feature vector by applying it to a residual vector quantizer; and applying the quantized result to a classifier model including a natural language processing model to output a classification label.
Owner:FOUND OF SOONGSIL UNIV IND COOP

System and method for knowledge-based audio-text modeling via automatic multimodal graph construction

Knowledge-based audio-text modeling via automatic multimodal graph construction is performed. An audio dataset is received, the audio dataset including clips of audio data, wherein each of the clips of the audio data is paired with corresponding metadata descriptive of the audio contents of the respective clip of the audio data. Graph nodes of interest are identified from a sematic network, the graph nodes being descriptive of semantics of the knowledge domain of the contents of the audio dataset. A large language model (LLM) is utilized for categorizing the metadata into the graph nodes and for inferring supplemental data for the graph nodes for which there is no metadata, producing an extracted knowledge graph. The extracted knowledge graph is validated utilizing the LLM to perform relation verification of edges between the graph nodes of the extracted knowledge graph, thereby mitigating hallucination effects in the categorizing and inferring of the supplemental data.
Owner:ROBERT BOSCH GMBH

Enterprise exhaustion method, device and equipment and storage medium

ActiveCN121352827AFinanceBiological modelsPersonalizationCall site
The invention provides an enterprise call-out method, device and equipment and a storage medium, and the method comprises the steps: obtaining the enterprise information of a to-be-called customer, converting the enterprise information into a label, and carrying out the matching of verbal skill contents from a historical verbal skill database according to the label, and obtaining a personalized verbal skill list; real-time voice recognition is carried out on an exhaustion call record collected on an exhaustion call site to obtain a real-time text, keyword retrieval is carried out on the real-time text, the real-time text is compared with the personalized verbal skill list, violation items and omission items are obtained respectively, and the violation items and the omission items are used for generating reminding information and sending the reminding information to exhaustion call personnel; after the exhaustion call process is finished, performing voice recognition on the exhaustion call record to obtain a full text, extracting key information from the full text and filling the key information into the report template to obtain an exhaustion call report; and acquiring the position information of the call-out site and the voiceprint characteristics of the call-out record, verifying the position information and the voiceprint characteristics, and marking the call-out report according to the verification result. The problem that in the prior art, the complete dispatch authenticity verification capability is weak is solved.
Owner:SHENGYE INFORMATION TECH SERVICE (SHENZHEN) CO LTD

Training Environmental Model for Premises Monitoring

A method of training a sound classification model for a premises monitoring system may include receiving audio data corresponding to a sound detected at a premises. The audio data may be provided to an active instance of a sound classification model, which may generate classification data indicating the sound is unrecognized. The audio data may be provided to a user device, and updated classification data indicating an identity of the sound may be received from the user device. The audio data and the updated classification data may be stored in a data bucket corresponding to the identity of the sound. When the data bucket contains a threshold quantity of user-classified audio data, a new instance of the sound classification model may be generated and retrained using the user-classified audio data. The active instance of the classification model may be replaced with the new instance of the sound classification model.
Owner:GROV LLC

Voice analyzer for interactive care system

A support interaction is guided in real time by generating from audio content featurized audio data that includes audio segments and audio features; generating in real time classification scores associated with certain audio segments; and displaying in real time the classifications scores and information associated with the corresponding audio segments.
Owner:LIVE CIRCLE INC

Interrogation record and audio and video recording linkage method and system

The invention provides an interrogation record and audio and video recording linkage method and system, and belongs to the technical field of information processing, and the method comprises the steps: obtaining interrogation record text data and segmented video recording data, and respectively labeling category labels; performing semantic extraction on the record text data to obtain semantic feature information, and converting the feature information into text semantic feature vectors; frame extraction is carried out on the video data, image features of each frame of the video data are extracted, similarity screening is carried out, and the image features are converted into key frame semantic feature vectors; constructing a cross-modal embedding space model, respectively encoding the text semantic feature vector and the key frame semantic feature vector of the category label through two encoders, then splicing and fusing to obtain a fused feature vector, converting and mapping the fused feature vector into a sharing layer, completing feature association of the record text data and the segmented video recording data, and obtaining the segmented video recording data. And subsequently, cross-modal search is carried out. According to the method, cross-modal information association is realized through association of the semantic feature vectors of the record and video recording information, and a basis is provided for subsequent rapid calling, research and judgment.
Owner:SHAANXI POLICE VOCATIONAL COLLEGE (SHAANXI POLITICAL & LEGAL MANAGEMENT CADRE COLLEGE)

Pet language translation method and system based on audio learning

The invention relates to a pet language translation method and system based on audio learning, and belongs to the technical field of animal training. The method comprises the following steps: constructing a pet standardized sound library; when the preset scene is triggered, playing the target sound signal; the target sound signal is a sound signal related to a preset scene in a pet standardized sound library; when it is detected that the first sound signal sent by the pet is matched with the sound signal in the pet standardized sound library, triggering a feedback operation corresponding to the matched sound signal; collecting pet sound in the current environment in real time, responding to the matching of the pet sound and the sound signal in the pet standardized sound library, and outputting a pet demand of the matched sound signal; the pet demand corresponds to a semantic translation result of the pet sound. By means of the mode, the training logic which can be stably recognized and can be copied and executed can be provided, accurate and reasonable pet language translation can be achieved, and the probability of mistranslation or wrong translation is reduced.
Owner:SHENZHEN KOLAMAMA TECH CO LTD

A computer assisted method for classifying digital audio files

A computer assisted method for classifying digital audio files based on features of a digital audio signal comprised in the file, comprising: storing the audio file in a digital memory; determining a portion (p) of drop of the audio file, for example the portion with the highest subjective loudness; classifying said audio file based on features of said drop.
Owner:WETWEAK SA

Intelligent music recommendation method based on emotion perception and acoustic characteristics

The invention discloses an intelligent music recommendation method based on emotional perception and acoustic features, and relates to the technical field of intelligent recommendation systems and emotional computation.The method comprises the steps that physiological signals, music acoustic features and historical behavior data of a user are collected; preprocessing the multi-source features and mapping the multi-source features to the same dimension to construct a fusion matrix; building a double-branch deep learning model, fusing features through an attention mechanism and training parameters; an evolutionary algorithm is adopted to optimize hyper-parameter screening optimal combination; generating a recommendation list matched with the real-time emotion and preference; and continuously collecting user interaction data, and regularly and incrementally training the dynamic update model. According to the method, emotion and behavior dual-drive recommendation is achieved by fusing physiological signals and acoustic features, emotion perception is accurate, recommended content fits the real-time mood, model optimization is efficient, recommendation precision and diversity are remarkably improved, and the music consumption experience of a user is greatly improved.
Owner:XIANGJIANG LAB

Audio Processing Engine Using Segmentation And Pruning

Techniques for diarization using embedding pruning are disclosed. A set of audio content segments and their associated tokens are accessed by a speaker enumeration module of a speech processing engine. The speaker enumeration module uses various pruning criteria to prune audio content segments from the set to result in a pruned set of audio content segments. The pruned set of audio content segments is analyzed using a clustering process to determine a number of speakers. The number of speakers is used in a second clustering process to identify speakers in the original set of audio content segments prior to pruning. A transcription of the original audio content with speaker labels is generated using the number of speakers identified for the pruned set of audio content segments.
Owner:ORACLE INT CORP

Archive management method based on artificial intelligence

The invention relates to the technical field of archive management, and discloses an archive management method based on artificial intelligence, comprising the following steps: collecting various types of archive data, including audio data, image data and text data; according to the method, through the steps of preprocessing, feature extraction, standardization and the like, archive data of different modalities are converted into unified multi-modal feature representation, data isomerism elimination and feature effective extraction are achieved, the multi-modal fusion method can fully utilize complementarity between data of different modalities, and the multi-modal fusion efficiency is improved. The overall expression ability of data is improved, a user can convert voice into a text and calculate semantic information of the text in a voice query mode, the semantic-level retrieval mode is more accurate and flexible than traditional keyword matching, a target file in a file database can be quickly retrieved based on input of a deep learning model, and the user experience is improved. And the retrieval efficiency and accuracy are improved.
Owner:WUHAN GOLDEN FILE TECH CO LTD

Engine abnormal sound detection method, device, equipment and medium

The invention discloses an engine abnormal sound detection method, device and equipment and a medium. The engine abnormal sound detection method comprises the following steps: acquiring actual noise audio data of a to-be-detected engine; processing the actual noise audio data to generate actual noise audio feature data; establishing a standard audio database; wherein the standard audio database comprises standard noise audio feature data under various combination conditions of various influence parameters; determining matched noise audio feature data according to the feature information of the to-be-detected engine and the standard noise audio feature data; according to similarity parameters of the matched noise audio feature data and the actual noise audio feature data, abnormal sound faults of the engine are recognized; and when the abnormal sound fault of the engine is identified, the abnormal sound fault type of the engine is diagnosed according to the actual noise audio feature data.
Owner:FAW JIEFANG AUTOMOTIVE CO

Automatic call categorization and screening

Implementations described herein relate to methods, systems, and computer-readable media to automatically answer a call. In some implementations, a method includes receiving a call from a caller device at a client device. The method further includes determining, based on an identifier associated with the call, whether the call matches auto answer criteria, and yin response to determining that the call matches the auto answer criteria, answering the call without user input and without alerting a user of the client device. The method further includes generating a call embedding for the call based on received audio of the call, comparing the call embedding with spam embeddings to determine whether the call is a spam call, and in response to determining that the call is a spam call, terminating the call.
Owner:GOOGLE LLC

Model training method, audio classification method, device, medium and program product

The application provides a model training method, an audio classification method, a device, a medium and a program product, and mainly relates to machine learning technology in the field of artificial intelligence. The training method comprises the following steps: obtaining a first audio, actual classification results and actual position encoding results of the first audio in multiple classification dimensions; inputting the first audio into a target neural network model to obtain predicted classification results and predicted position encoding results of the first audio in the multiple classification dimensions; obtaining a classification loss according to the actual classification results and the predicted classification results; fusing the actual position encoding results of the first audio in the multiple classification dimensions to obtain actual fusion results, and fusing the predicted position encoding results of the first audio in the multiple classification dimensions to obtain predicted fusion results; obtaining a position encoding loss according to the actual fusion results and the predicted fusion results; and training the target neural network model according to the classification loss and the position encoding loss, so that the classification accuracy can be improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Methods and apparatus to identify media

Methods, apparatus, systems and articles of manufacture are disclosed to identify media. An example method includes: in response to a query, generating an adjusted sample media fingerprint by applying an adjustment to a sample media fingerprint; comparing the adjusted sample media fingerprint to a reference media fingerprint; and in response to the adjusted sample media fingerprint matching the reference media fingerprint, transmitting information associated with the reference media fingerprint and the adjustment.
Owner:GRACENOTE INC

Techniques for audio track analysis to support audio personalization

Techniques for enabling personalization of audio tracks include selecting a portion of an audio track that is representative of the audio category, creating an audio sample from the portion of the audio track, playing the audio sample for a user, and adjusting, based on an input from the user while the audio sample is playing, a personalization setting for the user to be used when playing back audio from the audio category.
Owner:HARMAN INT IND INC

Music release disambiguation using multi-modal neural networks

ActiveUS12651166B2Neural architecturesNeural learning methodsMedicineMusic distribution
Methods and systems for disambiguating musical artist names are disclosed. Musical-artist-release records (MARRs) may be input to a multi-modal artificial neural network (ANN). Each MARR may be associated with a musical release of an artist, and may include a release ID and an artist ID, and release data in categories including music media content and metadata categories including sub-definitive musician name of the artist and release subcategories. All n-tuples of MARRs may be formed, and for each n-tuple, the ANN may be applied concurrently to each MARR to generate a release feature vector (RFV) that includes a set of sub-feature vectors, each characterizing a different category of release data. For each n-tuple, the ANN may be trained to cluster in a multi-dimensional RFV space RFVs of the same artist ID, and to separate RFVs of different artist IDs. The MARRs and their RFVs may be stored in a release database.
Owner:GRACENOTE INC

Music recommendation method and device and storage medium

The invention discloses a music recommendation method and device and a storage medium, and relates to the technical field of computers and the Internet. The method comprises the following steps: acquiring first position information, wherein the first position information is used for indicating a first position; obtaining candidate music matched with the first position information to obtain a first candidate set; based on the first position information, predicting the position of the terminal equipment at the second moment to obtain second position information; according to the first position information and the second position information, updating the first candidate set to obtain an updated first candidate set; at least one piece of recommended music is determined in the updated first candidate set, a recommended music set is obtained, and each piece of recommended music information is used for indicating one piece of recommended music. According to the method, the mode of recommending the music based on the current position information and the predicted position information is realized, so that the music recommended to the user can well adapt to the position change, and the music recommendation mode is enriched.
Owner:GUANGZHOU KUGOU COMP TECH CO LTD

Systems and methods for audio data enhancement

A method for training at least one machine learning model comprises receiving first audio data, receiving second audio data, and identifying, using labels associated with segments of the first audio data, non-target background audio segments in the first audio data.The method also includes identifying, using labels associated with segments of the second audio data, non-target background audio segments in the second audio data, and generating a first augmented training dataset by replacing the identified non-target background audio segments in the first audio data with the identified non-target background audio segments in the second audio data, generating a second augmented training dataset by replacing the identified non-target background audio segments in the second audio data with the identified non-target background audio segments in the first audio data, and training at least one machine learning model using the first augmented training dataset and the second augmented training dataset.
Owner:ROBERT BOSCH GMBH

Voice processing method, apparatus, electronic device, and computer program based on artificial intelligence

The present application provides an artificial intelligence-based audio processing method, device, electronic device, computer-readable storage medium, and computer program product, which relate to cloud technology and artificial intelligence technology. The method includes the steps of obtaining an audio clip of an audio scene, where the audio clip contains noise; performing an audio scene classification process based on the audio clip to obtain an audio scene type corresponding to the noise in the audio clip; determining a target audio processing mode corresponding to the audio scene type, and applying the target audio processing mode to the audio clip of the audio scene based on the degree of interference caused by the noise in the audio clip.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Music processing method, music processing device, program product and electronic device

The invention provides a music processing method, a music processing device, a program product and electronic equipment, and relates to the technical field of computers. The method comprises the following steps: displaying classification information of a first music set in response to a classification operation for the first music set; the classification information comprises one or more music subsets obtained by classifying music in the first music set through an artificial intelligence model; and playing music in a target subset in response to a playing operation for the target subset in the one or more music subsets. According to the invention, the operation process of music classification by the user is simplified, the dependence of the classification result on the label is reduced, and the diversity of the classification result is improved.
Owner:HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD

Method and apparatus for generating song list, electronic device, and storage medium

A method and an apparatus for generating a song list, an electronic device, a computer-readable storage medium, a computer program product and a computer program are provided. The method includes: acquiring candidate song library information, wherein the candidate song library information includes feature expressions of a candidate song, and the feature expressions represent song features in a plurality of dimensions; determining a similarity score of at least one candidate song according to the candidate song library information and a target feature expression, wherein the target feature expression is a feature expression of a seed song, and the similarity score represents the similarity between the candidate song and the seed song; and determining a target song based the similarity score of the candidate song, and generating a recommended song list based on the target song.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD +1

Music playing method and device, equipment and storage medium

The invention relates to a music playing method and device, equipment and a storage medium, and relates to the technical field of automobiles. The method comprises the steps that the music playing device obtains a feature vector of music to be played; furthermore, the music playing device determines target music according to the distance between the feature vector of the to-be-played music and the feature vector of each piece of preset music in the plurality of pieces of preset music, and the target music is the preset music which has the minimum distance with the feature vector of the to-be-played music in the plurality of pieces of preset music. Further, the music playing device determines a target playing sound effect according to the target music, and the target playing sound effect is a playing sound effect used when the target music is played or is a preset playing sound effect; and playing the to-be-played music based on the target playing sound effect. Therefore, the playing sound effect is automatically adjusted according to the to-be-played music, so that the playing effect of the music and the music experience of a user are improved.
Owner:DEEPAL AUTOMOBILE TECH CO LTD

Playlist generation method and apparatus, device, medium, product

The application relates to the technical field of music information retrieval, and discloses a playlist generation method and device, equipment, medium and product. The sorting method comprises the following steps: acquiring a music library knowledge graph, the knowledge graph is structured as a directed graph structure according to semantic correlation, and portrait labels of different songs in the music library are stored in a plurality of entity nodes in the directed graph structure; clustering the portrait labels of the entity nodes in the knowledge graph, generating a plurality of playlists corresponding to different portrait label combinations according to part of the portrait labels in the clustering result; acquiring portrait information of songs associated with the part of the portrait labels from the knowledge graph, determining portrait information of the playlists according to the acquired portrait information of the songs; and labeling the membership relationship between the playlists and the songs according to the similarity between the portrait information of the playlists and the portrait information of the songs. The application can realize automatic production of theme playlists and improve the quality of the theme playlists.
Owner:GUANGZHOU KUGOU COMP TECH CO LTD

A multimodal short video tag recommendation method integrating emotional information

The present invention discloses a multimodal short video tag recommendation method that integrates emotional information, belonging to the field of video processing technology. The method comprises the following steps: constructing a short video sample set; inputting the short video samples into an initial multimodal tag recommendation model based on a multi-head attention mechanism and an autoencoder, so that the model extracts features from the image, audio, and text of the short video samples to obtain content features and emotional features, and fuses these features using an attention network to obtain multiple candidate video tags; training the initial multimodal tag recommendation model with the desired video tag as the target and the difference in text features between the candidate video tags and the desired video tags as the loss to obtain a target multimodal tag recommendation model; and inputting the current short video into the target multimodal tag recommendation model to generate a target video tag. By integrating image features, audio features, and text features, the present invention can fully utilize multimodal information related to the video and effectively improve the quality of the generated video tags.
Owner:HUAZHONG UNIV OF SCI & TECH

Song list generation method and apparatus, and electronic device and storage medium

Provided in the embodiments of the present disclosure are a song list generation method and apparatus, and an electronic device, a computer-readable storage medium, a computer program product and a computer program. The method comprises: acquiring candidate song library information, wherein the candidate song library information comprises feature expressions of candidate songs, and the feature expressions represent song features in a plurality of dimensions; determining a similarity score of at least one candidate song according to the candidate song library information and a target feature expression, wherein the target feature expression is a feature expression of a seed song, and the similarity score represents the similarity between the candidate song and the seed song; and determining a target song on the basis of the similarity score of the candidate song, and generating a recommended song list on the basis of the target song. The similarity between the seed song and the candidate song is evaluated by using the feature expressions which represent the song features in the plurality of dimensions, and therefore a set of target songs which are more consistent with the seed song can be obtained, such that the recommended song list generated on the basis of the target songs has better consistency in the aspects of content, style, etc.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD +1