Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

97results about "Audio data clustering/classification" patented technology

A computer assisted method for classifying digital audio files

A computer assisted method for classifying digital audio files based on features of a digital audio signal comprised in the file, comprising: storing the audio file in a digital memory; determining a portion (p) of drop of the audio file, for example the portion with the highest subjective loudness; classifying said audio file based on features of said drop.
Owner:WETWEAK SA

Intelligent music recommendation method based on emotion perception and acoustic characteristics

The invention discloses an intelligent music recommendation method based on emotional perception and acoustic features, and relates to the technical field of intelligent recommendation systems and emotional computation.The method comprises the steps that physiological signals, music acoustic features and historical behavior data of a user are collected; preprocessing the multi-source features and mapping the multi-source features to the same dimension to construct a fusion matrix; building a double-branch deep learning model, fusing features through an attention mechanism and training parameters; an evolutionary algorithm is adopted to optimize hyper-parameter screening optimal combination; generating a recommendation list matched with the real-time emotion and preference; and continuously collecting user interaction data, and regularly and incrementally training the dynamic update model. According to the method, emotion and behavior dual-drive recommendation is achieved by fusing physiological signals and acoustic features, emotion perception is accurate, recommended content fits the real-time mood, model optimization is efficient, recommendation precision and diversity are remarkably improved, and the music consumption experience of a user is greatly improved.
Owner:XIANGJIANG LAB

Music release disambiguation using multi-modal neural networks

ActiveUS12651166B2Neural architecturesNeural learning methodsMedicineMusic distribution
Methods and systems for disambiguating musical artist names are disclosed. Musical-artist-release records (MARRs) may be input to a multi-modal artificial neural network (ANN). Each MARR may be associated with a musical release of an artist, and may include a release ID and an artist ID, and release data in categories including music media content and metadata categories including sub-definitive musician name of the artist and release subcategories. All n-tuples of MARRs may be formed, and for each n-tuple, the ANN may be applied concurrently to each MARR to generate a release feature vector (RFV) that includes a set of sub-feature vectors, each characterizing a different category of release data. For each n-tuple, the ANN may be trained to cluster in a multi-dimensional RFV space RFVs of the same artist ID, and to separate RFVs of different artist IDs. The MARRs and their RFVs may be stored in a release database.
Owner:GRACENOTE INC

Song list generation method and apparatus, and electronic device and storage medium

Provided in the embodiments of the present disclosure are a song list generation method and apparatus, and an electronic device, a computer-readable storage medium, a computer program product and a computer program. The method comprises: acquiring candidate song library information, wherein the candidate song library information comprises feature expressions of candidate songs, and the feature expressions represent song features in a plurality of dimensions; determining a similarity score of at least one candidate song according to the candidate song library information and a target feature expression, wherein the target feature expression is a feature expression of a seed song, and the similarity score represents the similarity between the candidate song and the seed song; and determining a target song on the basis of the similarity score of the candidate song, and generating a recommended song list on the basis of the target song. The similarity between the seed song and the candidate song is evaluated by using the feature expressions which represent the song features in the plurality of dimensions, and therefore a set of target songs which are more consistent with the seed song can be obtained, such that the recommended song list generated on the basis of the target songs has better consistency in the aspects of content, style, etc.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD +1

Cross-domain structural mapping in machine learning processing

A method of using a computing device executing to interrelate two or more corpuses of dissimilar data that includes receiving input data from each of two or more corpuses of dissimilar data. The computing device computes a pass for each of the input data into two or more encoder-decoder models. The computing device further obtains a prediction of an identity mapping for each of different domains of knowledge from each of the two or more encoder-decoder models. The computing device additionally computes a distribution distance metric as an output from each of a low-dimensional embedding vector representation from each of the two or more encoder-decoder models. The computing device still further computes a function based on each of the predictions from each of the two or more encoder-decoder models and the distribution distance metrics. The computing device additionally updates the two or more encoder-decoder models.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Systems and methods for machine learning-based classification of signal data signatures featuring using a multi-modal oracle

The disclosed systems and methods provide a novel technical solution via mechanisms for identifying which models are truly high-performing and the set of models that would provide the most accurate single prediction for a signal data signature (SDS). The disclosed systems and methods provides a computerized framework that can document the depictions of individual model performance. Moreover, the disclosed framework can identify all high performing models according to positive results, negative results, as well as generalized results. The framework can additionally operate to combine high performing models into a single predictive oracle to render a final prediction based on input from many models.
Owner:COVID COUGH INC

High-altitude projectile detection method and device based on multi-source signal recognition, and equipment

The application discloses a high-altitude projectile detection method and device based on multi-source signal identification, equipment and a storage medium, and the method comprises the following steps: acquiring a video and an audio, creating a vibe background model according to a first frame and detecting whether the audio is a projectile sound through a preset convolution model; acquiring a binary image of a current frame in real time through the vibe background model and updating the vibe background model; extracting the current frame and binary images before the current frame, subtracting all the binary images before the current frame from the binary image of the current frame to obtain a suspected projectile trajectory; judging whether the suspected projectile trajectory conforms to a projectile rule, and if yes, and the audio is a projectile sound, performing a projectile alarm. Through the method provided by the application, the vibe background model is used as a foreground checking algorithm, and the foreground noise caused by camera vibration and the like can be well inhibited, visual detection and audio recognition can better filter interference factors similar to the projectile trajectory, and the situation of detecting the projectile can be ensured to reduce false detection.
Owner:SHENZHEN INFINOVA

AI-based digital collection background music recommendation method

PendingCN121479009AMetadata audio data retrievalBiological modelsCosine similarityRegularization algorithm
The invention discloses an AI-based digital collection background music recommendation method, and relates to the technical field of AI. The method comprises the following steps: acquiring a key frame average image, a dominant tone vector and a style label of a digital collection, and candidate music samples and audio signals of a music library; extracting dynamic features of the digital collections to generate rhythm representation vectors; processing the music audio to generate a rhythm structure vector, and calculating the rhythm similarity between the rhythm structure vector and the rhythm structure vector through a dynamic time warping algorithm; generating a digital collection visual style vector and a music style vector, and calculating a style consistency score by using a cosine similarity algorithm; an auxiliary music feature vector is generated through the digital collection visual style vector, and a guide score is calculated in combination with the bottom layer audio features; and calculating a comprehensive score index according to the rhythm similarity, the style consistency score and the guide score. According to the invention, background music recommendation is carried out on the digital collections through the comprehensive scoring index.
Owner:湖北云雷信息技术有限公司

Systems and methods for data augmentation for multi-microphone signal processing

A method, computer program product, and computing system for receiving a signal from each microphone of a plurality of microphones, thereby defining a plurality of signals. One or more inter-microphone gain-based augmentations can be performed on the plurality of signals, thereby defining one or more inter-microphone gain-augmented signals. Performing one or more inter-microphone gain-based augmentations on the plurality of signals can include applying a gain level from a plurality of gain levels to the signal from each microphone. Applying a gain level from a plurality of gain levels to the signal from each microphone can include applying a gain level from a predefined range of gain levels to the signal from each microphone. Applying a gain level from a plurality of gain levels to the signal from each microphone can include applying a random gain level from a predefined range of gain levels to the signal from each microphone.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

A music search method and system based on big data

This invention discloses a music search method and system based on big data. It obtains sentiment characteristics by processing user historical playback data using a Natural Language Processing (NLP) model; acquires real-time semantic features based on real-time user input search context data using a BRET model; dynamically adjusts the matching weights of audio features, sentiment characteristics, and real-time semantic features using a reinforcement learning model; adjusts the feature weights based on the matching results to obtain fused features; inputs the fused features into an MTNN multi-task neural network, fuses the outputs of each branch through an attention mechanism to generate a music matching score, and generates a music recommendation list based on the music matching score. This improves the user search experience and platform conversion efficiency.
Owner:SICHUAN YUNSHUFUZHI EDUCATION TECH CO LTD

Multi-language automatic identification method and system

The invention relates to a multi-language automatic recognition method and system, and belongs to the technical field of language recognition, and the recognition method comprises the steps: receiving an original voice signal, and carrying out the preprocessing of the original voice signal, and obtaining a preprocessed voice frame sequence; synchronously extracting an acoustic feature vector, a vocal organ motion feature matrix and a rhythm feature vector from the voice frame sequence to form a feature triple; performing language family classification according to the acoustic feature vector and the vocal organ motion feature matrix, and outputting a candidate language family set; inputting the candidate language family set and the rhythm feature vector into a dialect clustering model, and outputting a refined dialect cluster tag; calculating an acoustic feature weight value, a vocal organ motion feature weight value and a rhythm feature weight value, and carrying out weighted operation on the feature triple to generate a weighted feature vector; and inputting the weighted feature vector and the refined dialect cluster label into a language decision model, and outputting a language recognition result containing a language label and a confidence value. According to the invention, the accuracy and robustness of multilingual recognition in a complex environment are improved.
Owner:BEIJING HIZHI TECH CO LTD

Automated call classification and screening

Implementations described herein relate to methods, systems, and computer-readable media for automatically answering calls. In some implementations, a method includes receiving, at a client device, a call from a calling device. The method also includes determining, based on an identifier related to the call, whether the call matches an automatic answer criteria, and in response to determining that the call matches the automatic answer criteria, answering the call without user input and without prompting a user of the client device. The method further includes generating a call embedding for the call based on audio received for the call, comparing the call embedding to a spam embedding to determine whether the call is a spam call, and in response to determining that the call is a spam call, terminating the call.
Owner:GOOGLE LLC

Latent spatial representation of audio signals for audio content-based capture

Methods and systems are provided for extracting features indicative of variations in pitch, timbre, decay, reverberation, and other psychoacoustic attributes from digital audio signals and training an artificial neural network model for generating, from the extracted features, a contextual latent space representation of the digital audio signals. Methods and systems are also provided for training an artificial neural network model for generating consistent latent space representations of digital audio signals, where the generated latent space representations are comparable for the purpose of determining psychoacoustic similarities between digital audio signals. Methods and systems are also provided for extracting features from digital audio signals and training an artificial neural network model for generating, from the extracted features, a latent space representation of the digital audio signals that is responsible for selecting salient attributes of the signals that are indicative of psychoacoustic differences between the signals.
Owner:DISTRIBUTED CREATION INC

Music recommendation method and system based on knowledge graph

PendingCN121958602AAchieve semantic connectivityImplement dynamic query capabilitiesMetadata audio data retrievalSpecial data processing applicationsPersonalizationKnowledge graph
The invention discloses a music recommendation method and system based on a knowledge graph. The content comprises multi-modal data collection, knowledge graph construction, multi-granularity clustering, fusion recall, adaptive perception optimization and intelligent recommendation. The invention relates to the technical field of music intelligent recommendation, in particular to a music recommendation method and system based on a knowledge graph, and the method comprises the steps: constructing a heterogeneous knowledge graph with a timestamp, accessing a time-aware graph self-attention network, attaching a multi-modal vector to a node, and combining a coding structure, content and time sequence attenuation; the clustering is softly injected into the topology, small-batch clustering is combined with neighbor indexes to realize large-scale soft distribution, and multi-granularity clustering is generated through modular community detection and theme labeling; clustering perception recall is adopted in retrieval, firstly, a user is mapped to a nearest cluster, then a long tail is expanded according to a high-weight path, long tail coverage and cold start capability is enhanced, more accurate personalized matching is achieved, and recommendation correlation and diversity are improved.
Owner:HUNAN INST OF INFORMATION TECH

Song recommendation method, electronic device, storage medium and computer program product

The invention discloses a song recommendation method and device, a storage medium and a computer program product. The method comprises the following steps: displaying a plurality of modified versions of a target song based on a controlled randomized display strategy; by taking the single display event as a unit, constructing a training sample and a corresponding sample tag based on the interactive behavior data for the multiple recompiled versions in the single display event and the context scene data when the single display event occurs; the audio representation of the target song, the context scene data and the reorganization style information of the reorganization version serve as input, the sample label serves as a supervision signal to train and optimize the sequencing model in the same song group to obtain a trained sequencing model in the same song group, and the trained sequencing model in the same song group is used for recommending the reorganization version of the target song. According to the method and the device, the scene data-based self-adaptive accurate recommendation is carried out on a plurality of recompiled versions of the same song.
Owner:TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD

A machine learning-based pain analysis and treatment regimen generation method and system

The application provides a pain analysis and treatment scheme generation method and system based on machine learning, which obtains pain information of a target patient, extracts feature information in the pain information, and enters the feature information into a feature information table; identifies a pain type of the target patient based on the feature information table; generates inquiry content based on the pain type; receives reply content of the target patient to the inquiry content; generates a pain treatment scheme of the target patient based on the reply content; obtains the feature information table by extracting pain information and mental and psychological evaluation results of the target patient, determines the pain type of the target patient according to the feature information table, and inquires the target patient according to the pain type and the mental and psychological state to supplement relevant information, and then generates a treatment scheme for the pain of the target patient, improves the accuracy of pain type identification based on the feature information table, and thus a more suitable treatment scheme can be recommended for the pain type of the target patient.
Owner:BEIJING SHIJITAN HOSPITAL CAPITAL MEDICAL UNIVERSITY

Method for collaborative knowledge base development

A use case knowledge base is collaboratively developed by receiving language input from a user, featurizing it into language elements, extracting predicate sets that are missing a predicate head or a predicate argument, querying users for input regarding the missing predicate information, and updating the knowledge base with predicate sets and other information provided by the users.
Owner:LIVE CIRCLE INC

Determining and tagging languages in audio files

Automatically detecting, tagging, and removing a human language stored in an audio file, including: training an application for detecting the human language using machine learning; loading each channel of the audio file into the trained application, wherein the audio file is an audio deliverable for motion picture and television; setting parameters and filtering each channel of the audio file to detect and tag the human language; and generating a list of timecodes and the corresponding human language detected.
Owner:SONY GROUP CORP +1

Audio-visual analytic for object rendering in capture

A system and method for the generation of automatic audio-visual analytics for object rendering in capture. One example provides a method of processing audiovisual content. The method includes receiving content including a plurality of audio frames and a plurality of video frames, classifying each of the plurality of audio frames into a plurality of audio classifications, and classifying each of the plurality of video frames into a plurality of video classifications. The method includes processing the plurality of audio frames based on the respective audio classifications and processing the plurality of video frames based on the respective video classifications. Each audio classification is processed with a different audio processing operation, and each video classification is processed with a different video processing operation. The method includes generating an audio / video representation of the content by merging the processed plurality of audio frames and the processed plurality of video frames.
Owner:DOLBY LABORATORIES LICENSING CORP

Song menu recommendation method and device, electronic equipment and storage medium

The invention discloses a song menu recommendation method and device, electronic equipment and a storage medium, and belongs to the technical field of recommendation. Based on interactive song menu features and interactive song features corresponding to an interactive song menu, fusion processing is performed to obtain interactive song menu song features corresponding to the interactive song menu; performing fusion processing on the basis of the candidate song menu features corresponding to the candidate song menu and the candidate song features to obtain candidate song menu song features corresponding to the candidate song menu; based on the interactive song menu song features and the candidate song menu song features, determining song menu song interest features for the candidate song menu, and based on the interactive song menu features and the candidate song menu features, determining the song menu interest features for the candidate song menu and song sequence features corresponding to the candidate song features and the interactive song; song interest features for the candidate song lists are determined; based on the song list song interest features, the song list interest features and the song interest features, the target song list is recommended for the target user from the candidate song lists, and the accuracy of song list recommendation is improved.
Owner:HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD

Enterprise due diligence method, apparatus, device, and storage medium

The application provides an enterprise due diligence method, device, equipment and storage medium, the method comprises the following steps: obtaining the enterprise information of the client to be due diligence and converting into a label, matching the content of the speech from the historical speech database according to the label, and obtaining the personalized speech list; real-time speech recognition is performed on the due diligence recording collected in the due diligence field to obtain real-time text, the real-time text is subjected to keyword retrieval and comparison with the personalized speech list, respectively obtaining the violation items and the missing items, which are used for generating reminder information and sending to the due diligence personnel; after the due diligence process is completed, the due diligence recording is subjected to speech recognition to obtain the full-amount text, the key information is extracted from the full-amount text and filled into the report template to obtain the due diligence report; the position information of the due diligence field and the voiceprint characteristics of the due diligence recording are collected, the position information and the voiceprint characteristics are verified, and the due diligence report is labeled according to the verification result. The application solves the problem of weak due diligence authenticity verification capability in the prior art.
Owner:SHENGYE INFORMATION TECH SERVICE (SHENZHEN) CO LTD

Systems and methods for data augmentation for multi-microphone signal processing

A method, computer program product, and computing system are provided for receiving signals from each of a plurality of microphones to define a plurality of signals. Harmonic distortion associated with at least one microphone can be determined. One or more harmonic distortion-based enhancements can be performed on the plurality of signals, at least in part, based on the harmonic distortion associated with the at least one microphone, thereby defining one or more harmonic distortion-based enhanced signals.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Sample data determination method, data processing method, apparatus, device, and medium

The application discloses a sample data determination method, a data processing method, a device, equipment and a medium, wherein the method comprises: acquiring an unlabeled data set; taking part of the unlabeled data from the unlabeled data set and labeling the part of the unlabeled data to obtain a first data set with labels corresponding to the part of the unlabeled data; training a type recognition model by using the first data set, the type recognition model being used for recognizing the unlabeled data to obtain a recognition result; acquiring corresponding remaining unlabeled data from the unlabeled data set; and determining a labeled sample data set based on the first data set, the type recognition model and the remaining unlabeled data. The application can save labor cost and time cost in the data labeling process, thereby improving the efficiency of data labeling.
Owner:HANGZHOU NETEASE ZHIQI TECH CO LTD

Apparatus and method for providing content

An apparatus for providing content includes a communication device that communicates with user equipment (UE) and a processor connected with the communication device. The processor analyzes a music database (DB) by interworking with the UE, extracts a driver emotion model based on the result of analyzing the music DB, determines an emotion determination model based on the result of analyzing the music DB, derives an emotional care correlation equation by means of a multi-regression analysis based on the result of analyzing the music DB, selects an emotional care solution depending on a contribution rate of the emotion determination model based on the driver emotion model using the emotional care correlation equation, and automatically play music content based on the emotional care solution.
Owner:HYUNDAI MOTOR CO LTD +2

A music category division method for multi-class multi-relation network

The application discloses a music category division method for a multi-class multi-relation network, belongs to the field of deep learning, and constructs a semantic double-path perception and topological similarity aggregation model for a multi-class multi-relation heterogeneous graph to divide music categories; the method comprises the following steps: obtaining music data of a current user and constructing a heterogeneous graph; constructing a topological feature similarity aggregation module, which recognizes nodes with similar structures by analyzing network connection modes based on topological similarity, and selects nodes with the highest topological similarity to aggregate features; constructing a semantic double-path perception aggregation module to aggregate complex semantic information in a differentiated learning multi-class multi-relation heterogeneous graph; and constructing a loss function to optimize a training model, so that music category division is realized. The application aims to provide users with a more efficient and accurate music classification experience, help users explore the richness and unique charm of music more deeply, and thus present a more splendid music world for music lovers.
Owner:SHANDONG UNIV OF SCI & TECH

Music classification method, music classification apparatus, electronic device, and storage medium

This application provides a music classification method, a music classification device, an electronic device, and a storage medium, belonging to the field of artificial intelligence technology. The method includes: acquiring sample audio data and sample lyrics data of sample music; extracting audio features from the sample audio data to obtain sample audio features; extracting lyric features from the sample lyrics data to obtain sample lyric features; constructing positive music sample pairs and negative music sample pairs based on the sample audio features and sample lyric features; training a neural network model based on the positive and negative music sample pairs to obtain a music classification model; acquiring target data for target music; extracting features from the target data to obtain target music features; scoring the target music according to its genre based on the music classification model and target music features to obtain genre score data; and determining the genre category of the target music based on the genre score data. This application can improve the accuracy of music classification.
Owner:PING AN TECH (SHENZHEN) CO LTD

Computer implemented method for song recommendation, and electronic device and storage medium

The present disclosure relates to the technical field of multimedia. Provided are a computer implemented method for song recommendation, and an electronic device and a storage medium. The method comprises: acquiring song-based real-time user interaction data and historical user usage data; on the basis of the real-time user interaction data and the historical user usage data, creating a target model, wherein the target model is a model that represents preferences of a user; acquiring a target music feature and a set of candidate songs similar to a currently played song, wherein the target music feature is a feature obtained by means of fusing multi-modal music features; and on the basis of the target model and the target music feature, creating a dynamic queue of the set of candidate songs, so as to obtain a song set to be played. The method enhances the diversity of data analysis factors for song recommendation, realizes the intelligent and dynamic adjustment of a music playing queue, provides a user with a highly personalized and dynamically responsive music playing experience, and improves the accuracy of played songs, thereby improving the personalized song listening experience of the user.
Owner:LINKPLAY TECHNOLOGY INC NANJING

A method for recognizing a piano piece, a computer device and a storage medium

The application relates to the field of music information retrieval, in particular to a piano piece identification method, a computer device and a storage medium. The method comprises the following steps: analyzing audio data in an audio to be identified, dividing the audio to be identified into a plurality of audio segments based on identified style transition nodes; generating a plurality of first audio fingerprints corresponding to the audio segments respectively; determining a corresponding performance style feature according to the audio data of each audio segment, and screening a general fingerprint library according to the performance style feature to determine a target fingerprint library corresponding to each audio segment; comparing the matching degree between the first audio fingerprint of each audio segment and the second audio fingerprint in the target fingerprint library corresponding to the audio segment to determine a plurality of candidate piano pieces corresponding to the audio segments; and determining at least one target piano piece from the plurality of candidate piano pieces based on the matching degree and / or the coincidence degree of the plurality of candidate piano pieces. The piano piece matching accuracy and efficiency are improved.
Owner:WUXIAN HONGYIN (CHONGQING) TECHNOLOGY CO LTD