Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

204results about "Audio data clustering/classification" patented technology

Adaptive sample selection for data item processing

Methods, systems, and apparatuses, including computer programs encoded on computer storage media, for receiving a query relating to a data item that includes multiple data item samples and processing the query and the data item to generate a response to the query. In particular, the described techniques include adaptively selecting a subset of the data item samples using a selection neural network conditioned on features of the data item samples and the query. Then processing the subset and query using a downstream task neural network to generate a response to the query. By adaptively selecting the subset of data item samples according to the query, the described techniques generate responses to queries that are more accurate and require less computation resources than would be the case using other techniques.
Owner:GOOGLE LLC

System and method for knowledge-based audio-text modeling via automatic multimodal graph construction

Knowledge-based audio-text modeling via automatic multimodal graph construction is performed. An audio dataset is received, the audio dataset including clips of audio data, wherein each of the clips of the audio data is paired with corresponding metadata descriptive of the audio contents of the respective clip of the audio data. Graph nodes of interest are identified from a sematic network, the graph nodes being descriptive of semantics of the knowledge domain of the contents of the audio dataset. A large language model (LLM) is utilized for categorizing the metadata into the graph nodes and for inferring supplemental data for the graph nodes for which there is no metadata, producing an extracted knowledge graph. The extracted knowledge graph is validated utilizing the LLM to perform relation verification of edges between the graph nodes of the extracted knowledge graph, thereby mitigating hallucination effects in the categorizing and inferring of the supplemental data.
Owner:ROBERT BOSCH GMBH

Enterprise exhaustion method, device and equipment and storage medium

ActiveCN121352827AFinanceBiological modelsPersonalizationCall site
The invention provides an enterprise call-out method, device and equipment and a storage medium, and the method comprises the steps: obtaining the enterprise information of a to-be-called customer, converting the enterprise information into a label, and carrying out the matching of verbal skill contents from a historical verbal skill database according to the label, and obtaining a personalized verbal skill list; real-time voice recognition is carried out on an exhaustion call record collected on an exhaustion call site to obtain a real-time text, keyword retrieval is carried out on the real-time text, the real-time text is compared with the personalized verbal skill list, violation items and omission items are obtained respectively, and the violation items and the omission items are used for generating reminding information and sending the reminding information to exhaustion call personnel; after the exhaustion call process is finished, performing voice recognition on the exhaustion call record to obtain a full text, extracting key information from the full text and filling the key information into the report template to obtain an exhaustion call report; and acquiring the position information of the call-out site and the voiceprint characteristics of the call-out record, verifying the position information and the voiceprint characteristics, and marking the call-out report according to the verification result. The problem that in the prior art, the complete dispatch authenticity verification capability is weak is solved.
Owner:SHENGYE INFORMATION TECH SERVICE (SHENZHEN) CO LTD

Pet language translation method and system based on audio learning

The invention relates to a pet language translation method and system based on audio learning, and belongs to the technical field of animal training. The method comprises the following steps: constructing a pet standardized sound library; when the preset scene is triggered, playing the target sound signal; the target sound signal is a sound signal related to a preset scene in a pet standardized sound library; when it is detected that the first sound signal sent by the pet is matched with the sound signal in the pet standardized sound library, triggering a feedback operation corresponding to the matched sound signal; collecting pet sound in the current environment in real time, responding to the matching of the pet sound and the sound signal in the pet standardized sound library, and outputting a pet demand of the matched sound signal; the pet demand corresponds to a semantic translation result of the pet sound. By means of the mode, the training logic which can be stably recognized and can be copied and executed can be provided, accurate and reasonable pet language translation can be achieved, and the probability of mistranslation or wrong translation is reduced.
Owner:SHENZHEN KOLAMAMA TECH CO LTD

A computer assisted method for classifying digital audio files

A computer assisted method for classifying digital audio files based on features of a digital audio signal comprised in the file, comprising: storing the audio file in a digital memory; determining a portion (p) of drop of the audio file, for example the portion with the highest subjective loudness; classifying said audio file based on features of said drop.
Owner:WETWEAK SA

Intelligent music recommendation method based on emotion perception and acoustic characteristics

The invention discloses an intelligent music recommendation method based on emotional perception and acoustic features, and relates to the technical field of intelligent recommendation systems and emotional computation.The method comprises the steps that physiological signals, music acoustic features and historical behavior data of a user are collected; preprocessing the multi-source features and mapping the multi-source features to the same dimension to construct a fusion matrix; building a double-branch deep learning model, fusing features through an attention mechanism and training parameters; an evolutionary algorithm is adopted to optimize hyper-parameter screening optimal combination; generating a recommendation list matched with the real-time emotion and preference; and continuously collecting user interaction data, and regularly and incrementally training the dynamic update model. According to the method, emotion and behavior dual-drive recommendation is achieved by fusing physiological signals and acoustic features, emotion perception is accurate, recommended content fits the real-time mood, model optimization is efficient, recommendation precision and diversity are remarkably improved, and the music consumption experience of a user is greatly improved.
Owner:XIANGJIANG LAB

Audio Processing Engine Using Segmentation And Pruning

Techniques for diarization using embedding pruning are disclosed. A set of audio content segments and their associated tokens are accessed by a speaker enumeration module of a speech processing engine. The speaker enumeration module uses various pruning criteria to prune audio content segments from the set to result in a pruned set of audio content segments. The pruned set of audio content segments is analyzed using a clustering process to determine a number of speakers. The number of speakers is used in a second clustering process to identify speakers in the original set of audio content segments prior to pruning. A transcription of the original audio content with speaker labels is generated using the number of speakers identified for the pruned set of audio content segments.
Owner:ORACLE INT CORP

Automatic call categorization and screening

Implementations described herein relate to methods, systems, and computer-readable media to automatically answer a call. In some implementations, a method includes receiving a call from a caller device at a client device. The method further includes determining, based on an identifier associated with the call, whether the call matches auto answer criteria, and yin response to determining that the call matches the auto answer criteria, answering the call without user input and without alerting a user of the client device. The method further includes generating a call embedding for the call based on received audio of the call, comparing the call embedding with spam embeddings to determine whether the call is a spam call, and in response to determining that the call is a spam call, terminating the call.
Owner:GOOGLE LLC

Model training method, audio classification method, device, medium and program product

The application provides a model training method, an audio classification method, a device, a medium and a program product, and mainly relates to machine learning technology in the field of artificial intelligence. The training method comprises the following steps: obtaining a first audio, actual classification results and actual position encoding results of the first audio in multiple classification dimensions; inputting the first audio into a target neural network model to obtain predicted classification results and predicted position encoding results of the first audio in the multiple classification dimensions; obtaining a classification loss according to the actual classification results and the predicted classification results; fusing the actual position encoding results of the first audio in the multiple classification dimensions to obtain actual fusion results, and fusing the predicted position encoding results of the first audio in the multiple classification dimensions to obtain predicted fusion results; obtaining a position encoding loss according to the actual fusion results and the predicted fusion results; and training the target neural network model according to the classification loss and the position encoding loss, so that the classification accuracy can be improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Techniques for audio track analysis to support audio personalization

Techniques for enabling personalization of audio tracks include selecting a portion of an audio track that is representative of the audio category, creating an audio sample from the portion of the audio track, playing the audio sample for a user, and adjusting, based on an input from the user while the audio sample is playing, a personalization setting for the user to be used when playing back audio from the audio category.
Owner:HARMAN INT IND INC

Music release disambiguation using multi-modal neural networks

ActiveUS12651166B2Neural architecturesNeural learning methodsMedicineMusic distribution
Methods and systems for disambiguating musical artist names are disclosed. Musical-artist-release records (MARRs) may be input to a multi-modal artificial neural network (ANN). Each MARR may be associated with a musical release of an artist, and may include a release ID and an artist ID, and release data in categories including music media content and metadata categories including sub-definitive musician name of the artist and release subcategories. All n-tuples of MARRs may be formed, and for each n-tuple, the ANN may be applied concurrently to each MARR to generate a release feature vector (RFV) that includes a set of sub-feature vectors, each characterizing a different category of release data. For each n-tuple, the ANN may be trained to cluster in a multi-dimensional RFV space RFVs of the same artist ID, and to separate RFVs of different artist IDs. The MARRs and their RFVs may be stored in a release database.
Owner:GRACENOTE INC

Method and apparatus for generating song list, electronic device, and storage medium

A method and an apparatus for generating a song list, an electronic device, a computer-readable storage medium, a computer program product and a computer program are provided. The method includes: acquiring candidate song library information, wherein the candidate song library information includes feature expressions of a candidate song, and the feature expressions represent song features in a plurality of dimensions; determining a similarity score of at least one candidate song according to the candidate song library information and a target feature expression, wherein the target feature expression is a feature expression of a seed song, and the similarity score represents the similarity between the candidate song and the seed song; and determining a target song based the similarity score of the candidate song, and generating a recommended song list based on the target song.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD +1

Playlist generation method and apparatus, device, medium, product

The application relates to the technical field of music information retrieval, and discloses a playlist generation method and device, equipment, medium and product. The sorting method comprises the following steps: acquiring a music library knowledge graph, the knowledge graph is structured as a directed graph structure according to semantic correlation, and portrait labels of different songs in the music library are stored in a plurality of entity nodes in the directed graph structure; clustering the portrait labels of the entity nodes in the knowledge graph, generating a plurality of playlists corresponding to different portrait label combinations according to part of the portrait labels in the clustering result; acquiring portrait information of songs associated with the part of the portrait labels from the knowledge graph, determining portrait information of the playlists according to the acquired portrait information of the songs; and labeling the membership relationship between the playlists and the songs according to the similarity between the portrait information of the playlists and the portrait information of the songs. The application can realize automatic production of theme playlists and improve the quality of the theme playlists.
Owner:GUANGZHOU KUGOU COMP TECH CO LTD

Song list generation method and apparatus, and electronic device and storage medium

Provided in the embodiments of the present disclosure are a song list generation method and apparatus, and an electronic device, a computer-readable storage medium, a computer program product and a computer program. The method comprises: acquiring candidate song library information, wherein the candidate song library information comprises feature expressions of candidate songs, and the feature expressions represent song features in a plurality of dimensions; determining a similarity score of at least one candidate song according to the candidate song library information and a target feature expression, wherein the target feature expression is a feature expression of a seed song, and the similarity score represents the similarity between the candidate song and the seed song; and determining a target song on the basis of the similarity score of the candidate song, and generating a recommended song list on the basis of the target song. The similarity between the seed song and the candidate song is evaluated by using the feature expressions which represent the song features in the plurality of dimensions, and therefore a set of target songs which are more consistent with the seed song can be obtained, such that the recommended song list generated on the basis of the target songs has better consistency in the aspects of content, style, etc.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD +1

Cross-domain structural mapping in machine learning processing

A method of using a computing device executing to interrelate two or more corpuses of dissimilar data that includes receiving input data from each of two or more corpuses of dissimilar data. The computing device computes a pass for each of the input data into two or more encoder-decoder models. The computing device further obtains a prediction of an identity mapping for each of different domains of knowledge from each of the two or more encoder-decoder models. The computing device additionally computes a distribution distance metric as an output from each of a low-dimensional embedding vector representation from each of the two or more encoder-decoder models. The computing device still further computes a function based on each of the predictions from each of the two or more encoder-decoder models and the distribution distance metrics. The computing device additionally updates the two or more encoder-decoder models.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Audio classification method and apparatus based on semi-supervised class incremental learning

The application relates to an audio classification method and device based on semi-supervised class incremental learning. The audio data to be classified is obtained, the audio data to be classified is input into a model obtained through task-by-task training on a data set corresponding to an ordered sequence task set based on a time sequence consistency regularization method and a loss function algorithm, and semi-supervised class incremental learning is performed to obtain the classification label of the audio data to be classified, thereby completing the classification of the audio. The time sequence consistency regularization method of the semi-supervised class incremental audio classification model is used to continuously update the parameters of the model, so that the matching degree between the model and the audio signal samples in the data set is higher and higher, thereby ensuring that the model is more and more accurate during classification. The audio samples generated in the loss function algorithm process of the semi-supervised class incremental audio classification model are stored, thereby effectively preventing the problem of catastrophic forgetting of the model during the learning process, and the stability of the model during audio classification is ensured.
Owner:NAT UNIV OF DEFENSE TECH

Multilingual speech and semantic intelligent translation method and system applied to exhibition scene

The invention discloses a multilingual speech semantic intelligent translation method and system applied to an exhibition scene, and belongs to the technical field of machine translation, and the method comprises the following steps: S1, obtaining a multi-person question judgment result; s2, if the multi-person questioning judgment result is multi-person questioning, audio identification information is obtained through analysis, and otherwise, the audio identification information is directly obtained through analysis; s3, obtaining each storage question keyword, each contrast question keyword and a key matching weighting factor corresponding to each question keyword; s4, obtaining a same-group evaluation result, if the same-group evaluation result is the same group, analyzing to obtain the comprehensive matching similarity of the storage groups, otherwise, analyzing the comprehensive matching similarity of each parallel storage group; s5, obtaining a comprehensive matching judgment result, if the comprehensive matching judgment result is unqualified, performing secondary refining processing to obtain a question and answer, and otherwise, directly obtaining the question and answer; and S6, voice broadcasting is carried out, and accurate separation and language recognition of voice sources of different questioning users are achieved.
Owner:ZHEJIANG HUIZHAN ELF TECHNOLOGY CO LTD

Systems and methods for machine learning-based classification of signal data signatures featuring using a multi-modal oracle

The disclosed systems and methods provide a novel technical solution via mechanisms for identifying which models are truly high-performing and the set of models that would provide the most accurate single prediction for a signal data signature (SDS). The disclosed systems and methods provides a computerized framework that can document the depictions of individual model performance. Moreover, the disclosed framework can identify all high performing models according to positive results, negative results, as well as generalized results. The framework can additionally operate to combine high performing models into a single predictive oracle to render a final prediction based on input from many models.
Owner:COVID COUGH INC

High-altitude projectile detection method and device based on multi-source signal recognition, and equipment

The application discloses a high-altitude projectile detection method and device based on multi-source signal identification, equipment and a storage medium, and the method comprises the following steps: acquiring a video and an audio, creating a vibe background model according to a first frame and detecting whether the audio is a projectile sound through a preset convolution model; acquiring a binary image of a current frame in real time through the vibe background model and updating the vibe background model; extracting the current frame and binary images before the current frame, subtracting all the binary images before the current frame from the binary image of the current frame to obtain a suspected projectile trajectory; judging whether the suspected projectile trajectory conforms to a projectile rule, and if yes, and the audio is a projectile sound, performing a projectile alarm. Through the method provided by the application, the vibe background model is used as a foreground checking algorithm, and the foreground noise caused by camera vibration and the like can be well inhibited, visual detection and audio recognition can better filter interference factors similar to the projectile trajectory, and the situation of detecting the projectile can be ensured to reduce false detection.
Owner:SHENZHEN INFINOVA

AI-based digital collection background music recommendation method

PendingCN121479009AMetadata audio data retrievalBiological modelsCosine similarityRegularization algorithm
The invention discloses an AI-based digital collection background music recommendation method, and relates to the technical field of AI. The method comprises the following steps: acquiring a key frame average image, a dominant tone vector and a style label of a digital collection, and candidate music samples and audio signals of a music library; extracting dynamic features of the digital collections to generate rhythm representation vectors; processing the music audio to generate a rhythm structure vector, and calculating the rhythm similarity between the rhythm structure vector and the rhythm structure vector through a dynamic time warping algorithm; generating a digital collection visual style vector and a music style vector, and calculating a style consistency score by using a cosine similarity algorithm; an auxiliary music feature vector is generated through the digital collection visual style vector, and a guide score is calculated in combination with the bottom layer audio features; and calculating a comprehensive score index according to the rhythm similarity, the style consistency score and the guide score. According to the invention, background music recommendation is carried out on the digital collections through the comprehensive scoring index.
Owner:湖北云雷信息技术有限公司

Characterization via homologizing disparate speech terminology

Aspects of the present disclosure are directed to methods and apparatuses involving characterization via homologizing disparate speech terminology. As may be implemented in accordance with one or more embodiments, audio processing circuitry is utilized to identify a respective language used for audio data sets. Homologizing circuitry is operable to homologize terms in the audio data sets for characterizing animals to which respective ones of the audio data sets are linked, by assessing and assigning terms in the respective audio data sets to respective homologized meanings based on the identified language for the audio data sets and an association between terms in the identified language for each audio data set and the homologized meaning. The homologized meanings may be in association with one of the animals to which the audio data set is linked, therein facilitating common characterizations of the animals utilizing disparate languages and terms.
Owner:WISCONSIN ALUMNI RES FOUND

Method for displaying information, method for searching for information and apparatus

A method for displaying information is provided. The method includes: receiving search information inputted by a user; acquiring multiple target search results corresponding to the search information, and multiple attribute tags; determining, for each of the attribute tags, at least one target search result under the attribute tag, where each of the target search results includes structured information; and displaying, on a search result display page, the multiple attribute tags and the at least one target search result under each of the attribute tags.
Owner:DOUYIN VISION CO LTD

Song search methods, devices, equipment, media, and products

This application discloses a song search method, apparatus, device, medium, and product. The method includes: obtaining the encoding information corresponding to a song submitted by a client; using a feature extraction model trained to convergence to extract a high-dimensional index vector representing deep semantic information at multiple scales of the song to be searched based on the encoding information; calculating the similarity between the high-dimensional index vector and high-dimensional index vectors representing deep semantic information at multiple scales of each candidate song extracted by the feature extraction model in a preset song feature library, obtaining a similarity sequence; filtering and determining target songs in the similarity sequence whose similarity values ​​exceed a preset threshold and are the most similar, and constructing a corresponding access link for the target song and pushing it to the client device. Through the above process, a song search service can be quickly, efficiently, and accurately realized, allowing users to find target songs similar to the song to be searched.
Owner:GUANGZHOU KUGOU COMP TECH CO LTD

Systems and methods for data augmentation for multi-microphone signal processing

A method, computer program product, and computing system for receiving a signal from each microphone of a plurality of microphones, thereby defining a plurality of signals. One or more inter-microphone gain-based augmentations can be performed on the plurality of signals, thereby defining one or more inter-microphone gain-augmented signals. Performing one or more inter-microphone gain-based augmentations on the plurality of signals can include applying a gain level from a plurality of gain levels to the signal from each microphone. Applying a gain level from a plurality of gain levels to the signal from each microphone can include applying a gain level from a predefined range of gain levels to the signal from each microphone. Applying a gain level from a plurality of gain levels to the signal from each microphone can include applying a random gain level from a predefined range of gain levels to the signal from each microphone.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

A music search method and system based on big data

This invention discloses a music search method and system based on big data. It obtains sentiment characteristics by processing user historical playback data using a Natural Language Processing (NLP) model; acquires real-time semantic features based on real-time user input search context data using a BRET model; dynamically adjusts the matching weights of audio features, sentiment characteristics, and real-time semantic features using a reinforcement learning model; adjusts the feature weights based on the matching results to obtain fused features; inputs the fused features into an MTNN multi-task neural network, fuses the outputs of each branch through an attention mechanism to generate a music matching score, and generates a music recommendation list based on the music matching score. This improves the user search experience and platform conversion efficiency.
Owner:SICHUAN YUNSHUFUZHI EDUCATION TECH CO LTD

Multi-language automatic identification method and system

The invention relates to a multi-language automatic recognition method and system, and belongs to the technical field of language recognition, and the recognition method comprises the steps: receiving an original voice signal, and carrying out the preprocessing of the original voice signal, and obtaining a preprocessed voice frame sequence; synchronously extracting an acoustic feature vector, a vocal organ motion feature matrix and a rhythm feature vector from the voice frame sequence to form a feature triple; performing language family classification according to the acoustic feature vector and the vocal organ motion feature matrix, and outputting a candidate language family set; inputting the candidate language family set and the rhythm feature vector into a dialect clustering model, and outputting a refined dialect cluster tag; calculating an acoustic feature weight value, a vocal organ motion feature weight value and a rhythm feature weight value, and carrying out weighted operation on the feature triple to generate a weighted feature vector; and inputting the weighted feature vector and the refined dialect cluster label into a language decision model, and outputting a language recognition result containing a language label and a confidence value. According to the invention, the accuracy and robustness of multilingual recognition in a complex environment are improved.
Owner:BEIJING HIZHI TECH CO LTD

Automated call classification and screening

Implementations described herein relate to methods, systems, and computer-readable media for automatically answering calls. In some implementations, a method includes receiving, at a client device, a call from a calling device. The method also includes determining, based on an identifier related to the call, whether the call matches an automatic answer criteria, and in response to determining that the call matches the automatic answer criteria, answering the call without user input and without prompting a user of the client device. The method further includes generating a call embedding for the call based on audio received for the call, comparing the call embedding to a spam embedding to determine whether the call is a spam call, and in response to determining that the call is a spam call, terminating the call.
Owner:GOOGLE LLC

Latent spatial representation of audio signals for audio content-based capture

Methods and systems are provided for extracting features indicative of variations in pitch, timbre, decay, reverberation, and other psychoacoustic attributes from digital audio signals and training an artificial neural network model for generating, from the extracted features, a contextual latent space representation of the digital audio signals. Methods and systems are also provided for training an artificial neural network model for generating consistent latent space representations of digital audio signals, where the generated latent space representations are comparable for the purpose of determining psychoacoustic similarities between digital audio signals. Methods and systems are also provided for extracting features from digital audio signals and training an artificial neural network model for generating, from the extracted features, a latent space representation of the digital audio signals that is responsible for selecting salient attributes of the signals that are indicative of psychoacoustic differences between the signals.
Owner:DISTRIBUTED CREATION INC

Multimedia equipment multi-period control method and system based on historical data analysis

The invention provides a multimedia equipment multi-period control method and system based on historical data analysis, and relates to the field of multimedia control. The method comprises the following steps: acquiring position information and weather forecast information of multimedia equipment; screening a reference playing record; determining an emotion regulation demand coefficient of the user; obtaining a multimedia content recommendation list; determining an emotion regulation confirmation coefficient of the user; determining whether the multimedia content recommendation list needs to be adjusted; obtaining an adjusted multimedia content recommendation list; and controlling the multimedia equipment to play. According to the method and the device, the emotion adjustment demand coefficient and the emotion adjustment confirmation coefficient can be determined by referring to the position information and the weather forecast information so as to judge whether the user needs to adjust the emotion at present, and if so, the playlist can be adjusted, the recommendation quantity of the injured music can be reduced, or the recommendation ranking of the injured music can be reduced. The possibility of continuously influencing the emotion of the user is reduced, and the emotion of the user can be adjusted.
Owner:WUXI FUTURE MIRROR DISPLAY TECH CO LTD