Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

98results about "Audio data indexing" patented technology

Enterprise exhaustion method, device and equipment and storage medium

ActiveCN121352827AFinanceBiological modelsPersonalizationCall site
The invention provides an enterprise call-out method, device and equipment and a storage medium, and the method comprises the steps: obtaining the enterprise information of a to-be-called customer, converting the enterprise information into a label, and carrying out the matching of verbal skill contents from a historical verbal skill database according to the label, and obtaining a personalized verbal skill list; real-time voice recognition is carried out on an exhaustion call record collected on an exhaustion call site to obtain a real-time text, keyword retrieval is carried out on the real-time text, the real-time text is compared with the personalized verbal skill list, violation items and omission items are obtained respectively, and the violation items and the omission items are used for generating reminding information and sending the reminding information to exhaustion call personnel; after the exhaustion call process is finished, performing voice recognition on the exhaustion call record to obtain a full text, extracting key information from the full text and filling the key information into the report template to obtain an exhaustion call report; and acquiring the position information of the call-out site and the voiceprint characteristics of the call-out record, verifying the position information and the voiceprint characteristics, and marking the call-out report according to the verification result. The problem that in the prior art, the complete dispatch authenticity verification capability is weak is solved.
Owner:SHENGYE INFORMATION TECH SERVICE (SHENZHEN) CO LTD

Short video copywriting tone automatic adjusting method driven by hierarchical rhythm mapping

The invention discloses a hierarchical rhythm mapping-driven short video copywriting mood automatic adjustment method, and relates to the technical field of video processing, and the method comprises the steps: 1, receiving a text character string and a language type identifier, and building an occupation column for bearing a tone mark, an accent mark and a duration mark at each level; 2, dividing each sentence into phrase segments based on the hierarchical index table, freezing boundaries by taking the phrase segments as units, presetting sentence end termination styles according to punctuations, determining kernel phrases according to semantic anchor points, initializing trends of the kernel phrases, and performing time sequence elastic alignment and hierarchical backfilling to obtain a sentence end termination pattern; and finally outputting a triple sequence which covers all syllables and is composed of tone marks, accent marks and duration marks as a target rhythm control sequence. And step 3, performing audio generation based on the target rhythm control sequence to obtain new dubbing. According to the method, the tone accuracy and expressive force of short video dubbing are improved, and the time and cost of manual adjustment are remarkably reduced.
Owner:CLOUD ATTACK NETWORK TECH HEBEI CO LTD

Electronic device stores tag information of content

An electronic device according to an embodiment comprises a memory, a display, and a processor operatively connected to the memory and the display, wherein the processor may be configured to: collect speech data; match the collected speech data with user information related to the collected speech data and store, in the memory, association information between the collected speech data and the user information; when generating content, detect speech data of the content input that is input during generation of the content; and when there is user information matching with the detected speech data in the memory, store the user information matching with the detected speech data of the content as tag information of the content.
Owner:SAMSUNG ELECTRONICS CO LTD

Audio gain output dynamic adjustment method and device, equipment, storage medium and computer program product

The invention relates to the technical field of audio processing, in particular to an audio gain output dynamic adjustment method and device, equipment, a storage medium and a computer program product. The method comprises the following steps: establishing a corresponding volume sequence based on initial audio data, and storing the volume sequence in a local database; based on the volume operation instruction of the user, updating the gain parameter, and recording boundary information corresponding to the volume operation instruction in a local database; determining a target playing range according to the boundary information, and generating a mapping relation between the input volume and the output volume; based on the mapping relation, processing the initial audio data by adopting a preset edge calculation algorithm to obtain an adjusted audio frame; and controlling the audio and video playing device to output the adjusted audio frame according to the updated gain parameter, thereby improving the audio playing quality of the audio and video playing device.
Owner:SHENZHEN JIUZHOU ELECTRIC

System and method for actionizing comments

A system and method for processing and actionizing structured and unstructured experience data is disclosed herein. In some embodiments, a system may include a natural language processing (NLP) engine configured to transform a data set into a plurality of concepts within a plurality of distinct contexts, and a data mining engine configured to process the relationships of the concepts and to identify associations and correlations in the data set. In some embodiments, the method may include the steps of receiving a data set, scanning the data set with an NLP engine to identify a plurality of concepts within a plurality of distinct contexts, and identifying patterns in the relationships between the plurality of concepts.
Owner:PRESS GANEY ASSOC LLC

Audio content segmentation and naming

Example implementations include dividing a textual transcript of digital audio content into a sequence of chunks, where the chunks are chronologically non-overlapping; determining annotations for each of the chunks, the annotations including at least one of: a title of the digital audio content, a description of the digital audio content, or one or more inferred segment titles of one or more previous segments of the digital audio content; providing, to a natural language model, a first chunk from the sequence of chunks, an associated annotation, and instructions to identify: a segment found in the first chunk, and a segment title of the segment; receiving, from the natural language model, an indication of the segment and the segment title; and storing the indication of the segment and the segment title as metadata associated with the digital audio content.
Owner:SPOTIFY

Incentivized electronic platform

A data structure embodied on a computer-readable medium is disclosed. The data structure may include database schema such as a structured query language (SQL) database. The database schema may include a registration schema that cooperates with a competition schema to award contestants engaged in a game of skill. The competition schema may encourage contestants to participate in games of skill related to songs, artists, and / or albums.
Owner:FAN LABEL LLC

Artificial neural network based search engine circuitry

Method (140, 200) and apparatus (120, 270) for characterizing digital content (124) using an artificial neural network (ANN) engine (122, 274). Computer data sets (126, 128, 130, 132, 134, 160, 202, 232, 242, 302, 332) from a library store (124) are processed to generate a corresponding sequence of multi-dimensional embedding vectors (162, 172, 182) in a latent space (170, 180). The embedding vectors are grouped into intervals or segments (166A, 168A, 228A) of the data sets based on movement metrics (164, 166, 168, 228) associated with the embedding vectors. A representative vector, RV (174A, 184B, 210, 276) is selected for each group. Thereafter, in response to a query input (272), selected intervals among the various computer data sets are identified and output based on a similarity measure (278) between the RVs and a search vector derived from the query input (150). Further embodiments provide a transformation model (322, 334) that transforms the embedding vectors and / or the RVs from a first latent space based on a first embedding model (304, 314, 332) to a different, second latent space based on a second embedding model (316, 338).
Owner:OBVIOUSFUTURE GMBH

Inspection report generation system, method and device, computer equipment and storage medium

The invention relates to an inspection report generation system, method and device, computer equipment and a storage medium. The system comprises a server, at least one terminal and recording devices corresponding to the terminals, the recording devices are used for recording dialogues between doctors and current patients in real time and sending recorded audio data streams to the corresponding terminals, and the server is used for establishing corresponding data transmission channels for the terminals and transmitting the data transmission channels to the terminals. The terminal is used for uploading the received audio data stream to the server in real time through the data transmission channel, and the server is used for performing character recognition on the audio data stream to obtain text information and generating an examination report corresponding to a current patient according to the text information. By adopting the method, the audio data can be ensured not to be disordered in the process from acquisition to transmission, a reliable data basis is provided for subsequent transcription and report generation, and the accuracy of a check report is improved.
Owner:BEIJING UNITED FAMILY HOSPITAL CO LTD

Data processing method and system based on artificial intelligence, storage device and storage medium

The invention discloses a data processing method and system based on artificial intelligence, a storage device and a storage medium, and relates to the technical field of storage. Inputting the data packet into a data processing model to generate a feature vector and a label so as to form structured metadata; globally unifying identifiers, organizing the feature vectors and the metadata into index structures, and mapping the feature vectors to an identifier list by each index structure; inputting the index structure into a data security verification model, and processing an abnormal data packet; inputting the index structure into a data processing model, and storing the index structure, the identifier of the original data item and the corresponding metadata in an index database; by accessing the index database to obtain the structured result and displaying the structured result to the user, the problem that the existing storage device is lack of intelligence is solved, and automatic classification, quick retrieval, secure encryption and efficient backup are realized.
Owner:PURPLELEC INC CO LTD

Micro-wave audio data intelligent processing and storage optimization system

The invention belongs to the technical field of artificial intelligence, and particularly relates to an intelligent processing and storage optimization system for micro-wave audio data, which comprises an audio data acquisition module, a real-time preprocessing module, an intelligent value evaluation module, a dynamic processing strategy generation module, a differential data processing module, a multi-stage storage optimization module and a system feedback and self-learning module. And dynamic and adaptive processing and storage optimization of the full life cycle of the audio data are realized through a data value driven integrated mechanism. The processing storage efficiency and the retrieval performance are effectively improved, and resource self-adaptive management and system intelligence are achieved.
Owner:GUANGZHOU YOUCAIHUA INFORMATION TECH CO LTD

Streaming music using supported services

An example technique includes a computing system storing media item identifiers of curated media items associated with one or more service providers. A media curating service aggregates the media item identifiers of curated media items. The example technique further involves receiving, from a media playback system, a first message comprising a service provider access identifier. The service provider access identifier is based on a user account of the media playback system registered to at least one service provider. Based on receiving the first message, the computing system determines media item identifiers of curated media items that are associated with the at least one service provider with which the user account of the media playback system is registered and causes the media playback system to play back the curated media items based on the determined media item identifiers of the curated media items.
Owner:SONOS INC

Audio playing control method and system based on identifier triggering, microphone, sound box equipment and storage medium

The invention discloses an audio playing control method and system based on identifier triggering, a microphone, sound box equipment and a storage medium. The method comprises the following steps: acquiring a non-contact identification signal; acquiring identification information corresponding to the non-contact identification signal; determining a corresponding target audio resource based on the identification information; and controlling an audio playing device to play the target audio resource. According to the method, the target audio resource is determined by using the non-contact identification signal, and the audio playing device is controlled to play the target audio resource, so that the sound box device can accurately position the audio content needing to be played according to the non-contact identification signal, a user does not need to carry out tedious manual operation or contact interaction, and the user experience is improved. The use convenience of the sound box equipment is improved, and the problems that the audio resource on-demand efficiency is low and even the audio resource on-demand operation cannot be completed due to the fact that the user is not familiar with the position of each function key on the sound box equipment or the screen operation logic are solved.
Owner:广东台德智联科技有限公司

AI-based digital collection background music recommendation method

PendingCN121479009AMetadata audio data retrievalBiological modelsCosine similarityRegularization algorithm
The invention discloses an AI-based digital collection background music recommendation method, and relates to the technical field of AI. The method comprises the following steps: acquiring a key frame average image, a dominant tone vector and a style label of a digital collection, and candidate music samples and audio signals of a music library; extracting dynamic features of the digital collections to generate rhythm representation vectors; processing the music audio to generate a rhythm structure vector, and calculating the rhythm similarity between the rhythm structure vector and the rhythm structure vector through a dynamic time warping algorithm; generating a digital collection visual style vector and a music style vector, and calculating a style consistency score by using a cosine similarity algorithm; an auxiliary music feature vector is generated through the digital collection visual style vector, and a guide score is calculated in combination with the bottom layer audio features; and calculating a comprehensive score index according to the rhythm similarity, the style consistency score and the guide score. According to the invention, background music recommendation is carried out on the digital collections through the comprehensive scoring index.
Owner:湖北云雷信息技术有限公司

Characterization via homologizing disparate speech terminology

Aspects of the present disclosure are directed to methods and apparatuses involving characterization via homologizing disparate speech terminology. As may be implemented in accordance with one or more embodiments, audio processing circuitry is utilized to identify a respective language used for audio data sets. Homologizing circuitry is operable to homologize terms in the audio data sets for characterizing animals to which respective ones of the audio data sets are linked, by assessing and assigning terms in the respective audio data sets to respective homologized meanings based on the identified language for the audio data sets and an association between terms in the identified language for each audio data set and the homologized meaning. The homologized meanings may be in association with one of the animals to which the audio data set is linked, therein facilitating common characterizations of the animals utilizing disparate languages and terms.
Owner:WISCONSIN ALUMNI RES FOUND

A Personalized Reconstruction Method of Head-Related Transfer Function Based on Spatial Orientation Fusion and Frequency Channel Fusion

This invention discloses a personalized reconstruction method for the head-related transfer function (HRTF) based on spatial orientation fusion and frequency channel fusion, aiming to solve the technical problem of how to quickly and accurately obtain a comprehensive personalized HRTF of a subject from a small amount of measured orientation data. It includes the following steps: preprocessing HRTF data in the CIPIC database; rearranging the three-dimensional amplitude spectra of all orientations at all pitch angles after preprocessing to obtain a two-dimensional amplitude spectrum of spatial orientation-frequency channels; retaining the amplitude values ​​of all frequencies in the spatial orientation portion of the two-dimensional amplitude spectrum, and setting the amplitude values ​​of other orientations to 0, to obtain the input dataset; establishing a neural network structure for personalized HRTF reconstruction; and inputting the preprocessed data into the neural network structure for training to form a neural network model for personalized HRTF reconstruction. The model of this invention has low complexity, exhibits good performance in terms of mean logarithmic spectral distortion and root mean square error, and has a short training time.
Owner:ZHENGZHOU UNIV

Song search methods, devices, equipment, media, and products

This application discloses a song search method, apparatus, device, medium, and product. The method includes: obtaining the encoding information corresponding to a song submitted by a client; using a feature extraction model trained to convergence to extract a high-dimensional index vector representing deep semantic information at multiple scales of the song to be searched based on the encoding information; calculating the similarity between the high-dimensional index vector and high-dimensional index vectors representing deep semantic information at multiple scales of each candidate song extracted by the feature extraction model in a preset song feature library, obtaining a similarity sequence; filtering and determining target songs in the similarity sequence whose similarity values ​​exceed a preset threshold and are the most similar, and constructing a corresponding access link for the target song and pushing it to the client device. Through the above process, a song search service can be quickly, efficiently, and accurately realized, allowing users to find target songs similar to the song to be searched.
Owner:GUANGZHOU KUGOU COMP TECH CO LTD

Automated call classification and screening

Implementations described herein relate to methods, systems, and computer-readable media for automatically answering calls. In some implementations, a method includes receiving, at a client device, a call from a calling device. The method also includes determining, based on an identifier related to the call, whether the call matches an automatic answer criteria, and in response to determining that the call matches the automatic answer criteria, answering the call without user input and without prompting a user of the client device. The method further includes generating a call embedding for the call based on audio received for the call, comparing the call embedding to a spam embedding to determine whether the call is a spam call, and in response to determining that the call is a spam call, terminating the call.
Owner:GOOGLE LLC

Latent spatial representation of audio signals for audio content-based capture

Methods and systems are provided for extracting features indicative of variations in pitch, timbre, decay, reverberation, and other psychoacoustic attributes from digital audio signals and training an artificial neural network model for generating, from the extracted features, a contextual latent space representation of the digital audio signals. Methods and systems are also provided for training an artificial neural network model for generating consistent latent space representations of digital audio signals, where the generated latent space representations are comparable for the purpose of determining psychoacoustic similarities between digital audio signals. Methods and systems are also provided for extracting features from digital audio signals and training an artificial neural network model for generating, from the extracted features, a latent space representation of the digital audio signals that is responsible for selecting salient attributes of the signals that are indicative of psychoacoustic differences between the signals.
Owner:DISTRIBUTED CREATION INC

Automated audio description system and method

An audio description system includes a memory and a processor. The memory stores source media comprising frames positioned within the source media according to a time index. The processor is configured to generate, using an image-to-text model, a textual description of each frame; identify intervals within the time index, each interval encompassing one or more positions of one or more frames; identify placement periods within the time index, each placement period being temporally proximal to an interval; generate a summary description based on at least one textual description of at least one frame positioned within a selected interval temporally proximal to a placement period; and associate the summary description with the placement period.
Owner:3PLAY MEDIA

A song management method, device, equipment, storage medium and program product

The application discloses a song management method, device and equipment, a storage medium and a program product. The method comprises the following steps: obtaining data source information of a song to be added to a cross-platform playlist; the song in the cross-platform playlist comes from at least two multimedia platforms; obtaining identification information and song index information of the song according to the data source information; generating the cross-platform playlist containing the identification information of the song, and the identification information of the song is associated with the song index information. Thus, the storage and maintenance of the cross-platform playlist are realized, the unified management of songs with different sources in the same platform is facilitated, the switching between different multimedia platforms for song management is avoided, and the efficiency of music resource management is improved.
Owner:GUANGZHOU KUGOU COMP TECH CO LTD

Methods and systems for processing audio signals containing speech data

Methods and systems for processing audio signals containing speech data are disclosed. Biometric data associated with at least one speaker are extracted from an audio input. A correspondence is determined between the extracted biometric data and stored biometric data associated with a consenting user profile, where a consenting user profile is a user profile indicates consent to store biometric data. If no correspondence is determined, the speech data is discarded, optionally after having been processed.
Owner:SOAPBOX LABS LTD

Song cover synthesis, performing method and device, equipment, medium and product

The application discloses a song cover synthesis, execution method and device, equipment, medium and product. The synthesis method comprises the following steps: obtaining a target song specified by a user through a graphical user interface; obtaining ranking data of a plurality of singing singers matched with a singing feature in the target song from a server; displaying individual controls corresponding to each singing singer according to the ranking data, and determining a corresponding target singing singer by touch control; and in response to a touch event acting on any individual control, obtaining a cover song in which the target singing singer sings the target song from the server according to the individual control. The application makes the production of the cover song more convenient, improves the production efficiency of the cover song, enriches the music auxiliary creation form, and improves the user experience.
Owner:GUANGZHOU KUGOU COMP TECH CO LTD

Methods and systems for updating a facial digital image in an electronic document

Embodiments of the present disclosure relates to a server, a user device, and methods for updating a facial digital image in an electronic document. A method comprises the steps of receiving, from a user, a selection of a facial digital image update option at a user interface of a software application executed on a user device. The method further comprises the steps of receiving a new facial digital image upon the selection of the facial digital image update option. The method also comprises the steps of transmitting a digital certificate generation request to a server. The method also comprises the steps of receiving a digital certificate comprising the new facial digital image and a new SOD and updating the electronic document with the new facial digital image and the new SOD associated with the new facial digital image.
Owner:THALES DIS FRANCE SA

Metaverse Personalized Digital Singer Generation System and Method Thereof

A metaverse personalized digital singer generation system and a method thereof. In the system, the server-end device receives a user voice, store the user voice as a personalized voice, capture an image of a user face to generate a facial image, generate a personalized digital singer displayed in a virtual scene through a 3D imaging technology, and convert the personalized voice into voice feature vectors, and use the voice feature vectors and the personalized voice as training data, input the training data to a generative AI model to train a generative pre-training model having the personal characteristics. When the user selects an original song for singing, the original song and the user singing voice and the prompt are inputted to the generative pre-training model, the remixed song matching a style of the original song is outputted, a vocal coaching is generated and displayed based on prompt.
Owner:SQ TECH (SHANGHAI) CORP +1

Systems and methods for partitioning search indexes for improved efficiency in identifying media segments

Systems and methods for identifying a media segment of audio or video content are described. The video segment is identified by deriving data from media content and comparing said data to a reference database in order to identify said video segment. Embodiments of the invention improve the speed and accuracy of the media identification process by advantageously partitioning the indexes in subdivisions where high value reference information is separated from the bulk information, for example.
Owner:VIZIO INSCAPE TECH LLC

A laparoscopic surgery video retrieval and visualization method and system

The present application relates to a kind of laparoscopic surgery video retrieval and visualization method and system, which is achieved by "video analysis-data organization-multi-mode retrieval" efficient surgery video retrieval.In the video analysis stage, the deep learning model is combined with the hidden Markov model of medical knowledge constraint, automatically analyzes the video content and divides the standardization process unit, reduces the consumption of computing resources and training time.In the data organization stage, the feature entity bidirectional mapping is proposed, which converts the unstructured video data into a hierarchical data structure with medical semantics, facilitating storage and retrieval.In the multi-mode retrieval stage, three retrieval modes of time axis orientation, instrument orientation and process unit orientation are designed, combined with multi-track display and medical semantic visualization coding to realize dynamic visualization.Through the interactive feedback of multi-mode retrieval method and time sequence feature dynamic visualization, a professional retrieval experience that meets the training and research needs of medical personnel is provided.
Owner:HUNAN UNIV

Index system of brand appliance coding library and voice control method thereof

The application provides an index system of a brand appliance coding library and a voice control method thereof, and belongs to the field of household appliance remote controllers. A three-layer index system is set for an infrared coding library module, and then a target brand appliance is identified in response to a target pairing voice input by a user, the target brand appliance is matched and associated with corresponding infrared numbers based on a third index library module, a second index library module, a first index library module and the infrared coding library module, a target infrared coding association set is obtained, and then the target infrared coding association set is sent to the target brand appliance in sequence in response to a target operation voice input by the user for the target brand appliance, so that one-key household appliance pairing and control are realized, the pairing process is simplified, and user operation is facilitated.
Owner:SHENZHEN CHAORAN TECH CO LTD

Song matching method and apparatus, device, medium, and product

The application discloses a song matching method and device, equipment, medium and product, and the method comprises the following steps: obtaining the encoding information corresponding to the audio data of the to-be-matched karaoke melody submitted by a client; using a feature extraction model trained to a convergence state to extract a high-dimensional index vector representing the multi-scale deep semantic information of the to-be-matched karaoke melody according to the encoding information; calculating the similarity between the high-dimensional index vector and the high-dimensional index vector representing the multi-scale deep semantic information of each melody segment extracted by the feature extraction model in the melody feature library, and screening out a target melody segment satisfying a preset condition; and pushing a target song containing the target melody segment in a song library to a client device. Through the above process, the song search service can be quickly and efficiently realized, and the user can find a target song similar to the to-be-matched karaoke melody.
Owner:GUANGZHOU KUGOU COMP TECH CO LTD