Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

167results about "Audio data indexing" patented technology

Media content management

A system and method for media content management include creating, via a digital vault, a container file comprising media content submitted by a first user and content metadata; verifying, via the digital vault, a completeness of the content metadata associated with the media content in the container file; classifying, via the digital vault, the container file based on the completeness of the media content; capturing, via the digital vault, event metadata when a second user gains access to the container file, the event metadata comprising at least one of identification of the second user, an activation timestamp, a duration of access, portions of the container file accessed, and changes to the container file; and enabling a private communication channel between parties affiliated with the media content to permit messaging among the parties affiliated with the media content via the private communication channel.
Owner:TUNEGO INC

Enterprise exhaustion method, device and equipment and storage medium

ActiveCN121352827AFinanceBiological modelsPersonalizationCall site
The invention provides an enterprise call-out method, device and equipment and a storage medium, and the method comprises the steps: obtaining the enterprise information of a to-be-called customer, converting the enterprise information into a label, and carrying out the matching of verbal skill contents from a historical verbal skill database according to the label, and obtaining a personalized verbal skill list; real-time voice recognition is carried out on an exhaustion call record collected on an exhaustion call site to obtain a real-time text, keyword retrieval is carried out on the real-time text, the real-time text is compared with the personalized verbal skill list, violation items and omission items are obtained respectively, and the violation items and the omission items are used for generating reminding information and sending the reminding information to exhaustion call personnel; after the exhaustion call process is finished, performing voice recognition on the exhaustion call record to obtain a full text, extracting key information from the full text and filling the key information into the report template to obtain an exhaustion call report; and acquiring the position information of the call-out site and the voiceprint characteristics of the call-out record, verifying the position information and the voiceprint characteristics, and marking the call-out report according to the verification result. The problem that in the prior art, the complete dispatch authenticity verification capability is weak is solved.
Owner:SHENGYE INFORMATION TECH SERVICE (SHENZHEN) CO LTD

Short video copywriting tone automatic adjusting method driven by hierarchical rhythm mapping

The invention discloses a hierarchical rhythm mapping-driven short video copywriting mood automatic adjustment method, and relates to the technical field of video processing, and the method comprises the steps: 1, receiving a text character string and a language type identifier, and building an occupation column for bearing a tone mark, an accent mark and a duration mark at each level; 2, dividing each sentence into phrase segments based on the hierarchical index table, freezing boundaries by taking the phrase segments as units, presetting sentence end termination styles according to punctuations, determining kernel phrases according to semantic anchor points, initializing trends of the kernel phrases, and performing time sequence elastic alignment and hierarchical backfilling to obtain a sentence end termination pattern; and finally outputting a triple sequence which covers all syllables and is composed of tone marks, accent marks and duration marks as a target rhythm control sequence. And step 3, performing audio generation based on the target rhythm control sequence to obtain new dubbing. According to the method, the tone accuracy and expressive force of short video dubbing are improved, and the time and cost of manual adjustment are remarkably reduced.
Owner:CLOUD ATTACK NETWORK TECH HEBEI CO LTD

Multi-language interactive learning system based on speech recognition

The invention relates to the field of voice signal processing, in particular to a multi-language interactive learning system based on voice recognition, which is characterized in that sound waves and mouth shape images are synchronously acquired and discretized by the system, and cross-modal coding is formed after alignment; secondly, the code is injected into a micro-ring photon reserve network through phase modulation to unfold time sequence characteristics, the code is mapped into a quaternion graph to be embedded, and segment boundaries are extracted through a diffusion-pulse coupling method; then, according to a fixed field sequence, encapsulating the quaternion graph embedding and segmentation data into an object mark prompt, inputting the object mark prompt into a low-rank adaptive language model, and generating semantic segmentation data and text transcription; and finally, the adaptive learning module implements sparse gradient updating on the low-rank weight and pulse network by using an integer fractal hash exclusive-or difference mask, and adopts support-query element learning for synchronous iteration after user clarification. According to the system, low-power-consumption, high-precision and second-level accent self-adaptive multi-language voice interaction is realized on the end side.
Owner:SICHUAN COLLEGE OF ARCHITECTURAL TECH

Automatic call categorization and screening

Implementations described herein relate to methods, systems, and computer-readable media to automatically answer a call. In some implementations, a method includes receiving a call from a caller device at a client device. The method further includes determining, based on an identifier associated with the call, whether the call matches auto answer criteria, and yin response to determining that the call matches the auto answer criteria, answering the call without user input and without alerting a user of the client device. The method further includes generating a call embedding for the call based on received audio of the call, comparing the call embedding with spam embeddings to determine whether the call is a spam call, and in response to determining that the call is a spam call, terminating the call.
Owner:GOOGLE LLC

Electronic device stores tag information of content

An electronic device according to an embodiment comprises a memory, a display, and a processor operatively connected to the memory and the display, wherein the processor may be configured to: collect speech data; match the collected speech data with user information related to the collected speech data and store, in the memory, association information between the collected speech data and the user information; when generating content, detect speech data of the content input that is input during generation of the content; and when there is user information matching with the detected speech data in the memory, store the user information matching with the detected speech data of the content as tag information of the content.
Owner:SAMSUNG ELECTRONICS CO LTD

Music data processing method and device

The invention provides a music data processing method and device, electronic equipment, a non-instantaneous computer readable storage medium and a computer program product, and belongs to the technical field of music data processing. The method comprises the following steps: acquiring music data; marking the music data as a predetermined music structure according to a first preset rule; wherein the preset music structure comprises at least one of music segments, music sentences, music knots and phonetic types; marking the preset music structures as preset attributes according to a second preset rule, and marking the relationship between the preset music structures to form marked music data; according to the method, the annotated music data is stored in the corresponding data table of the database, the problem that annotation of the music data cannot meet training data requirements in the prior art is solved, and the quality and efficiency of annotation of the music data are improved, so that the training data requirements are better met, and data support with annotations is provided for subsequent task research.
Owner:BEIJING INSTITUTE FOR GENERAL ARTIFICIAL INTELLIGENCE

Digital video production systems and methods

Described herein is a computer implemented method. The method includes displaying, on a display, a scene timeline including a time-ordered sequence of scene previews, each scene preview corresponding to a scene of a video production and having a display width that provides a visual indication of a duration of that scene. The method further includes displaying a canvas including a first visual element that is associated with the first scene, and in response to detecting selection of the first visual element from the canvas, causing a first visual element timing indicator to be displayed. The first visual element timing indicator is aligned with the scene timeline based on a first visual element start time and a first visual element end time.
Owner:CANVA PTY LTD

Audio gain output dynamic adjustment method and device, equipment, storage medium and computer program product

The invention relates to the technical field of audio processing, in particular to an audio gain output dynamic adjustment method and device, equipment, a storage medium and a computer program product. The method comprises the following steps: establishing a corresponding volume sequence based on initial audio data, and storing the volume sequence in a local database; based on the volume operation instruction of the user, updating the gain parameter, and recording boundary information corresponding to the volume operation instruction in a local database; determining a target playing range according to the boundary information, and generating a mapping relation between the input volume and the output volume; based on the mapping relation, processing the initial audio data by adopting a preset edge calculation algorithm to obtain an adjusted audio frame; and controlling the audio and video playing device to output the adjusted audio frame according to the updated gain parameter, thereby improving the audio playing quality of the audio and video playing device.
Owner:SHENZHEN JIUZHOU ELECTRIC

System and method for actionizing comments

A system and method for processing and actionizing structured and unstructured experience data is disclosed herein. In some embodiments, a system may include a natural language processing (NLP) engine configured to transform a data set into a plurality of concepts within a plurality of distinct contexts, and a data mining engine configured to process the relationships of the concepts and to identify associations and correlations in the data set. In some embodiments, the method may include the steps of receiving a data set, scanning the data set with an NLP engine to identify a plurality of concepts within a plurality of distinct contexts, and identifying patterns in the relationships between the plurality of concepts.
Owner:PRESS GANEY ASSOC LLC

Audio content segmentation and naming

Example implementations include dividing a textual transcript of digital audio content into a sequence of chunks, where the chunks are chronologically non-overlapping; determining annotations for each of the chunks, the annotations including at least one of: a title of the digital audio content, a description of the digital audio content, or one or more inferred segment titles of one or more previous segments of the digital audio content; providing, to a natural language model, a first chunk from the sequence of chunks, an associated annotation, and instructions to identify: a segment found in the first chunk, and a segment title of the segment; receiving, from the natural language model, an indication of the segment and the segment title; and storing the indication of the segment and the segment title as metadata associated with the digital audio content.
Owner:SPOTIFY

Incentivized electronic platform

A data structure embodied on a computer-readable medium is disclosed. The data structure may include database schema such as a structured query language (SQL) database. The database schema may include a registration schema that cooperates with a competition schema to award contestants engaged in a game of skill. The competition schema may encourage contestants to participate in games of skill related to songs, artists, and / or albums.
Owner:FAN LABEL LLC

Artificial neural network based search engine circuitry

Method (140, 200) and apparatus (120, 270) for characterizing digital content (124) using an artificial neural network (ANN) engine (122, 274). Computer data sets (126, 128, 130, 132, 134, 160, 202, 232, 242, 302, 332) from a library store (124) are processed to generate a corresponding sequence of multi-dimensional embedding vectors (162, 172, 182) in a latent space (170, 180). The embedding vectors are grouped into intervals or segments (166A, 168A, 228A) of the data sets based on movement metrics (164, 166, 168, 228) associated with the embedding vectors. A representative vector, RV (174A, 184B, 210, 276) is selected for each group. Thereafter, in response to a query input (272), selected intervals among the various computer data sets are identified and output based on a similarity measure (278) between the RVs and a search vector derived from the query input (150). Further embodiments provide a transformation model (322, 334) that transforms the embedding vectors and / or the RVs from a first latent space based on a first embedding model (304, 314, 332) to a different, second latent space based on a second embedding model (316, 338).
Owner:OBVIOUSFUTURE GMBH

Inspection report generation system, method and device, computer equipment and storage medium

The invention relates to an inspection report generation system, method and device, computer equipment and a storage medium. The system comprises a server, at least one terminal and recording devices corresponding to the terminals, the recording devices are used for recording dialogues between doctors and current patients in real time and sending recorded audio data streams to the corresponding terminals, and the server is used for establishing corresponding data transmission channels for the terminals and transmitting the data transmission channels to the terminals. The terminal is used for uploading the received audio data stream to the server in real time through the data transmission channel, and the server is used for performing character recognition on the audio data stream to obtain text information and generating an examination report corresponding to a current patient according to the text information. By adopting the method, the audio data can be ensured not to be disordered in the process from acquisition to transmission, a reliable data basis is provided for subsequent transcription and report generation, and the accuracy of a check report is improved.
Owner:BEIJING UNITED FAMILY HOSPITAL CO LTD

Data processing method and system based on artificial intelligence, storage device and storage medium

The invention discloses a data processing method and system based on artificial intelligence, a storage device and a storage medium, and relates to the technical field of storage. Inputting the data packet into a data processing model to generate a feature vector and a label so as to form structured metadata; globally unifying identifiers, organizing the feature vectors and the metadata into index structures, and mapping the feature vectors to an identifier list by each index structure; inputting the index structure into a data security verification model, and processing an abnormal data packet; inputting the index structure into a data processing model, and storing the index structure, the identifier of the original data item and the corresponding metadata in an index database; by accessing the index database to obtain the structured result and displaying the structured result to the user, the problem that the existing storage device is lack of intelligence is solved, and automatic classification, quick retrieval, secure encryption and efficient backup are realized.
Owner:PURPLELEC INC CO LTD

Micro-wave audio data intelligent processing and storage optimization system

The invention belongs to the technical field of artificial intelligence, and particularly relates to an intelligent processing and storage optimization system for micro-wave audio data, which comprises an audio data acquisition module, a real-time preprocessing module, an intelligent value evaluation module, a dynamic processing strategy generation module, a differential data processing module, a multi-stage storage optimization module and a system feedback and self-learning module. And dynamic and adaptive processing and storage optimization of the full life cycle of the audio data are realized through a data value driven integrated mechanism. The processing storage efficiency and the retrieval performance are effectively improved, and resource self-adaptive management and system intelligence are achieved.
Owner:GUANGZHOU YOUCAIHUA INFORMATION TECH CO LTD

Automatic call categorization and screening

Implementations described herein relate to methods, systems, and computer-readable media to automatically answer a call. In some implementations, a method includes receiving a call from a caller device at a client device. The method further includes determining, based on an identifier associated with the call, whether the call matches auto answer criteria, and yin response to determining that the call matches the auto answer criteria, answering the call without user input and without alerting a user of the client device. The method further includes generating a call embedding for the call based on received audio of the call, comparing the call embedding with spam embeddings to determine whether the call is a spam call, and in response to determining that the call is a spam call, terminating the call.
Owner:GOOGLE LLC

Music education method and system based on artificial intelligence

The invention discloses a music education method and system based on artificial intelligence, and relates to the technical field of music education, and the music education method based on artificial intelligence comprises the following steps: obtaining music learning data of a user; based on a preset classification mode, performing cloud service matching on the music learning data of the user to obtain a plurality of original storage positions; artificial intelligence is applied to the field of music education, learning data of a user is analyzed through the deep learning model, making of a learning scheme and resource recommendation are achieved, and the learning efficiency and effect are improved. According to the method, historical music learning data is utilized to construct a deep learning model for feature analysis and pattern recognition, so that the learning demand and the learning scheme of a user are predicted, and more accurate learning guidance is provided for the user.
Owner:HAINAN NORMAL UNIV

Streaming music using supported services

An example technique includes a computing system storing media item identifiers of curated media items associated with one or more service providers. A media curating service aggregates the media item identifiers of curated media items. The example technique further involves receiving, from a media playback system, a first message comprising a service provider access identifier. The service provider access identifier is based on a user account of the media playback system registered to at least one service provider. Based on receiving the first message, the computing system determines media item identifiers of curated media items that are associated with the at least one service provider with which the user account of the media playback system is registered and causes the media playback system to play back the curated media items based on the determined media item identifiers of the curated media items.
Owner:SONOS INC

Audio playing control method and system based on identifier triggering, microphone, sound box equipment and storage medium

The invention discloses an audio playing control method and system based on identifier triggering, a microphone, sound box equipment and a storage medium. The method comprises the following steps: acquiring a non-contact identification signal; acquiring identification information corresponding to the non-contact identification signal; determining a corresponding target audio resource based on the identification information; and controlling an audio playing device to play the target audio resource. According to the method, the target audio resource is determined by using the non-contact identification signal, and the audio playing device is controlled to play the target audio resource, so that the sound box device can accurately position the audio content needing to be played according to the non-contact identification signal, a user does not need to carry out tedious manual operation or contact interaction, and the user experience is improved. The use convenience of the sound box equipment is improved, and the problems that the audio resource on-demand efficiency is low and even the audio resource on-demand operation cannot be completed due to the fact that the user is not familiar with the position of each function key on the sound box equipment or the screen operation logic are solved.
Owner:广东台德智联科技有限公司

A large model-based voice dialogue retrieval method, device and medium

The application discloses a large model-based voice dialogue retrieval method, device and medium, and belongs to the technical field of voice retrieval. The method comprises the following steps: constructing a to-be-retrieved database and an inverted index database, wherein the former stores a text field, and the latter converts a phonetic alphabet. A user is guided to input a retrieval voice by using a preset dialogue template. The voice is converted into a phonetic alphabet retrieval information by using a voice recognition large model, and then a fuzzy matching algorithm is used to compare the inverted index and filter high matching degree fields. If the matching fields are many, a filtering voice is generated to further inquire; if the matching fields are few or none, the user is prompted to re-input or correct the input; and if the matching degree is 1, the retrieval field is directly determined. According to the field, the result in the to-be-retrieved database is found, and a unique direct reply is given; if the result is not unique, a missing field is identified and inquired, so that accurate retrieval is realized. The whole process is guided by a dialogue template. The application realizes the effects of flexible retrieval mode and low cold start cost to a certain extent by the above method.
Owner:SHANDONG SYNTHESIS ELECTRONICS TECH

AI-based digital collection background music recommendation method

PendingCN121479009AMetadata audio data retrievalBiological modelsCosine similarityRegularization algorithm
The invention discloses an AI-based digital collection background music recommendation method, and relates to the technical field of AI. The method comprises the following steps: acquiring a key frame average image, a dominant tone vector and a style label of a digital collection, and candidate music samples and audio signals of a music library; extracting dynamic features of the digital collections to generate rhythm representation vectors; processing the music audio to generate a rhythm structure vector, and calculating the rhythm similarity between the rhythm structure vector and the rhythm structure vector through a dynamic time warping algorithm; generating a digital collection visual style vector and a music style vector, and calculating a style consistency score by using a cosine similarity algorithm; an auxiliary music feature vector is generated through the digital collection visual style vector, and a guide score is calculated in combination with the bottom layer audio features; and calculating a comprehensive score index according to the rhythm similarity, the style consistency score and the guide score. According to the invention, background music recommendation is carried out on the digital collections through the comprehensive scoring index.
Owner:湖北云雷信息技术有限公司

Characterization via homologizing disparate speech terminology

Aspects of the present disclosure are directed to methods and apparatuses involving characterization via homologizing disparate speech terminology. As may be implemented in accordance with one or more embodiments, audio processing circuitry is utilized to identify a respective language used for audio data sets. Homologizing circuitry is operable to homologize terms in the audio data sets for characterizing animals to which respective ones of the audio data sets are linked, by assessing and assigning terms in the respective audio data sets to respective homologized meanings based on the identified language for the audio data sets and an association between terms in the identified language for each audio data set and the homologized meaning. The homologized meanings may be in association with one of the animals to which the audio data set is linked, therein facilitating common characterizations of the animals utilizing disparate languages and terms.
Owner:WISCONSIN ALUMNI RES FOUND

A Personalized Reconstruction Method of Head-Related Transfer Function Based on Spatial Orientation Fusion and Frequency Channel Fusion

This invention discloses a personalized reconstruction method for the head-related transfer function (HRTF) based on spatial orientation fusion and frequency channel fusion, aiming to solve the technical problem of how to quickly and accurately obtain a comprehensive personalized HRTF of a subject from a small amount of measured orientation data. It includes the following steps: preprocessing HRTF data in the CIPIC database; rearranging the three-dimensional amplitude spectra of all orientations at all pitch angles after preprocessing to obtain a two-dimensional amplitude spectrum of spatial orientation-frequency channels; retaining the amplitude values ​​of all frequencies in the spatial orientation portion of the two-dimensional amplitude spectrum, and setting the amplitude values ​​of other orientations to 0, to obtain the input dataset; establishing a neural network structure for personalized HRTF reconstruction; and inputting the preprocessed data into the neural network structure for training to form a neural network model for personalized HRTF reconstruction. The model of this invention has low complexity, exhibits good performance in terms of mean logarithmic spectral distortion and root mean square error, and has a short training time.
Owner:ZHENGZHOU UNIV

Song search methods, devices, equipment, media, and products

This application discloses a song search method, apparatus, device, medium, and product. The method includes: obtaining the encoding information corresponding to a song submitted by a client; using a feature extraction model trained to convergence to extract a high-dimensional index vector representing deep semantic information at multiple scales of the song to be searched based on the encoding information; calculating the similarity between the high-dimensional index vector and high-dimensional index vectors representing deep semantic information at multiple scales of each candidate song extracted by the feature extraction model in a preset song feature library, obtaining a similarity sequence; filtering and determining target songs in the similarity sequence whose similarity values ​​exceed a preset threshold and are the most similar, and constructing a corresponding access link for the target song and pushing it to the client device. Through the above process, a song search service can be quickly, efficiently, and accurately realized, allowing users to find target songs similar to the song to be searched.
Owner:GUANGZHOU KUGOU COMP TECH CO LTD

Method for controlling range extender, and related device

Provided are a method for controlling a range extender and an apparatus therefor, and an electronic device, a vehicle, a non-transitory computer-readable storage medium, a computer program product and a computer program. The method comprises: calling, from a pre-established range extender working condition database, a corresponding working condition, a sound pressure level of which is less than or equal to a sound pressure level margin; using the corresponding working condition as a correction target working condition of a range extender; then, performing correction according to an acquired sound pressure level of ambient noise and a sound pressure level of the range extender under each working condition in the pre-established range extender working condition database by means of a correction sound pressure function, so as to obtain a corrected sound pressure level; and finally, comparing a sound pressure level corresponding to the correction target working condition of the range extender with the corrected sound pressure level, and adjusting the correction target working condition of the range extender according to a comparison result..
Owner:BEIJING CO WHEELS TECH CO LTD

Automated call classification and screening

Implementations described herein relate to methods, systems, and computer-readable media for automatically answering calls. In some implementations, a method includes receiving, at a client device, a call from a calling device. The method also includes determining, based on an identifier related to the call, whether the call matches an automatic answer criteria, and in response to determining that the call matches the automatic answer criteria, answering the call without user input and without prompting a user of the client device. The method further includes generating a call embedding for the call based on audio received for the call, comparing the call embedding to a spam embedding to determine whether the call is a spam call, and in response to determining that the call is a spam call, terminating the call.
Owner:GOOGLE LLC

Latent spatial representation of audio signals for audio content-based capture

Methods and systems are provided for extracting features indicative of variations in pitch, timbre, decay, reverberation, and other psychoacoustic attributes from digital audio signals and training an artificial neural network model for generating, from the extracted features, a contextual latent space representation of the digital audio signals. Methods and systems are also provided for training an artificial neural network model for generating consistent latent space representations of digital audio signals, where the generated latent space representations are comparable for the purpose of determining psychoacoustic similarities between digital audio signals. Methods and systems are also provided for extracting features from digital audio signals and training an artificial neural network model for generating, from the extracted features, a latent space representation of the digital audio signals that is responsible for selecting salient attributes of the signals that are indicative of psychoacoustic differences between the signals.
Owner:DISTRIBUTED CREATION INC

Matching audio fingerprints

Methods, apparatus, systems and articles of manufacture are disclosed to select reference sub-fingerprints for comparison to query sub-fingerprints based on a determination that a query sub-fingerprint is a match with a reference sub-fingerprint, generate a count vector that stores total counts of matches between the query sub-fingerprints and different subsets of the reference sub-fingerprints, each of the different subsets being aligned to the query sub-fingerprints at a different offset from a reference point, each of the different offsets being mapped by the count vector to a different total count, calculate a maximum count among the total counts, a median of the total counts, and a difference between the maximum count and the median of the total counts, and classify the reference sub-fingerprints as a match with the query sub-fingerprints based on the difference between the maximum count in the count vector and the median.
Owner:GRACENOTE INC