Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

112results about "Audio data querying" patented technology

Multi-mode collaborative board game entertainment method and system based on artificial intelligence

The invention provides a multi-mode collaborative board game entertainment method and system based on artificial intelligence, and relates to the technical field of artificial intelligence. The network server is innovated in a mode of AI voice recognition, an AI host model and a multi-modal generation model, game automation is achieved, and multi-modal feedback of characters, voice and dynamic scenes is provided; wherein the AI host can automatically determine the current progress condition of a game and generate an optional plot branch text in combination with a player instruction, the collaborative architecture multi-modal AI model generates multi-modal data matched with the plot branch text according to the plot branch text, and the multi-modal data has consistency and accuracy, enriches the output form of the board game entertainment system, and improves the performance of the board game entertainment system. And flexible and changeable multi-mode game guidance is provided.
Owner:SHANGHAI ZHICHEN TECHNOLOGY CO LTD

Music database retrieval method and system based on feature extraction

The invention discloses a music database retrieval method and system based on feature extraction, and relates to the technical field of data analysis, and the method comprises the steps: carrying out the preprocessing of an audio signal of a query track input by a user, dividing the audio into equal-length frame sequences, and generating three-channel spectrogram tensors, including a Mel spectrogram, a logarithmic amplitude spectrogram and a CQT spectrogram; extracting multi-modal features of the equal-length frame sequence, dividing melody motivation segments according to pitch change and stability of the audio signal, performing differential coding, generating a melody differential sequence, and modeling the melody differential sequence into a melody topological graph; and calculating the structural similarity between the melody topological graph and a database melody graph, screening a similar melody candidate set P, mapping a melody curve into a topological manifold through a topological data analysis method, extracting persistent homology features, and converting the persistent homology features into a persistent bar graph. According to the method, the retrieval precision and the matching credibility of the complex melody in the music database are remarkably improved.
Owner:BODA COLLEGE OF JILIN NORMAL UNIV

Methods and apparatus to identify media

Methods, apparatus, systems and articles of manufacture are disclosed to identify media. An example method includes: in response to a query, generating an adjusted sample media fingerprint by applying an adjustment to a sample media fingerprint; comparing the adjusted sample media fingerprint to a reference media fingerprint; and in response to the adjusted sample media fingerprint matching the reference media fingerprint, transmitting information associated with the reference media fingerprint and the adjustment.
Owner:GRACENOTE INC

Digital video production systems and methods

Described herein is a computer implemented method. The method includes displaying, on a display, a scene timeline including a time-ordered sequence of scene previews, each scene preview corresponding to a scene of a video production and having a display width that provides a visual indication of a duration of that scene. The method further includes displaying a canvas including a first visual element that is associated with the first scene, and in response to detecting selection of the first visual element from the canvas, causing a first visual element timing indicator to be displayed. The first visual element timing indicator is aligned with the scene timeline based on a first visual element start time and a first visual element end time.
Owner:CANVA PTY LTD

Bundled search processing method and apparatus, computer device, and storage medium

The application relates to a bundle search processing method and device, computer equipment and a storage medium. The method comprises the following steps: performing bundle search on a word table based on a current to-be-matched object and a current bundle width to obtain a current search result, wherein the current search result comprises a plurality of target words which are matched successfully with the current to-be-matched object and the number of which matches the current bundle width; performing width attenuation on the current bundle width according to a bundle width attenuation mode matched by the number of executed bundle searches to obtain a bundle width for next bundle search; taking each target word as a to-be-matched object for next bundle search, and iteratively performing bundle search until the bundle width obtained by attenuation is equal to a bundle width threshold value, and taking the search result at this time as a target to-be-matched object, and performing bundle search on the word table based on the bundle width threshold value to obtain a target search result. By adopting the method, the search efficiency can be improved while ensuring the search effect in an ultra-long sequence scene.
Owner:ZHAOLIAN CONSUMER FINANCE CO LTD

Audio playing control method and system based on identifier triggering, microphone, sound box equipment and storage medium

The invention discloses an audio playing control method and system based on identifier triggering, a microphone, sound box equipment and a storage medium. The method comprises the following steps: acquiring a non-contact identification signal; acquiring identification information corresponding to the non-contact identification signal; determining a corresponding target audio resource based on the identification information; and controlling an audio playing device to play the target audio resource. According to the method, the target audio resource is determined by using the non-contact identification signal, and the audio playing device is controlled to play the target audio resource, so that the sound box device can accurately position the audio content needing to be played according to the non-contact identification signal, a user does not need to carry out tedious manual operation or contact interaction, and the user experience is improved. The use convenience of the sound box equipment is improved, and the problems that the audio resource on-demand efficiency is low and even the audio resource on-demand operation cannot be completed due to the fact that the user is not familiar with the position of each function key on the sound box equipment or the screen operation logic are solved.
Owner:广东台德智联科技有限公司

Voice query QOS based on client-computed content metadata

A method includes receiving an automated speech recognition (ASR) request from a user device that includes a speech input captured by the user device and content metadata associated with the speech input. The content metadata is generated by the user device. The method also includes determining a priority score for the ASR request based on the content metadata associated with the speech input and caching the ASR request in a pre-processing backlog of pending ASR requests each having a corresponding priority score. The pending ASR requests in the pre-processing backlog are ranked in order of the priority scores. The method also includes providing, from the pre-processing backlog, one or more of the pending ASR requests to a backend-side ASR module, wherein pending ASR requests associated with higher priority scores are processed before pending ASR requests associated with lower priority scores.
Owner:GOOGLE LLC

system

We provide the system. [Solution] Means for obtaining user image information, A means for analyzing the aforementioned image information to infer the atmosphere and emotional state, means for generating musical information based on the aforementioned atmosphere and emotional state, Means for transmitting the generated music information to the user's terminal device, A system that includes this.
Owner:SOFTBANK GROUP CORP

Multilingual speech and semantic intelligent translation method and system applied to exhibition scene

The invention discloses a multilingual speech semantic intelligent translation method and system applied to an exhibition scene, and belongs to the technical field of machine translation, and the method comprises the following steps: S1, obtaining a multi-person question judgment result; s2, if the multi-person questioning judgment result is multi-person questioning, audio identification information is obtained through analysis, and otherwise, the audio identification information is directly obtained through analysis; s3, obtaining each storage question keyword, each contrast question keyword and a key matching weighting factor corresponding to each question keyword; s4, obtaining a same-group evaluation result, if the same-group evaluation result is the same group, analyzing to obtain the comprehensive matching similarity of the storage groups, otherwise, analyzing the comprehensive matching similarity of each parallel storage group; s5, obtaining a comprehensive matching judgment result, if the comprehensive matching judgment result is unqualified, performing secondary refining processing to obtain a question and answer, and otherwise, directly obtaining the question and answer; and S6, voice broadcasting is carried out, and accurate separation and language recognition of voice sources of different questioning users are achieved.
Owner:ZHEJIANG HUIZHAN ELF TECHNOLOGY CO LTD

Song searching method, device, searching system and computer readable storage medium

The application discloses a song search method and device, a search system and a computer readable storage medium, and belongs to the technical field of search. The embodiment receives a song search request, wherein the song search request comprises song search information; searches a first song from a first storage space corresponding to a first search engine based on the song search information through the first search engine, wherein the first storage space stores song information of all songs in a music platform; searches a second song from a second storage space corresponding to a second search engine based on the song search information through the second search engine, wherein the second storage space stores song information of songs meeting preset conditions in the music platform; and generates a search result corresponding to the song search request based on the first song and the second song, so that the stability of the search system and the accuracy of the search result can be improved.
Owner:HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD

Compensating for Time Scale Differences to Facilitate Audio Identification

A method includes receiving, by a computing system, an audio signal, where the audio signal defines a segment of media content over time. The method also includes establishing by the computing system, based on the received audio signal, a normalized query frequency-domain representation of the received audio signal. The method further includes matching, by the computing system, the normalized query frequency-domain representation of the received audio signal with a correspondingly normalized reference frequency-domain representation of a reference audio signal having an associated identity. The method additionally includes based on the matching, determining by the computing system that an identity of the received audio signal is the associated identity of the reference audio signal.
Owner:THE NIELSEN CO (US) LLC

Method and device for identifying media

A method and apparatus for identifying media. Methods, apparatus, systems, and articles of manufacture to identify media are disclosed. An example method includes, in response to a query, generating an adjusted sample media fingerprint by applying an adjustment to the sample media fingerprint; comparing the adjusted sample media fingerprint with a reference media fingerprint; and in response to matching of the adjusted sample media fingerprint and the reference media fingerprint, sending information associated with the reference media fingerprint and the adjustment.
Owner:GRACENOTE INC

Information processing device, information processing method, and information processing program

To provide an information processing device, an information processing method, and an information processing program that are capable of increasing the convenience of users.SOLUTION: An information processing device includes a specification unit, an acquisition unit, and an image processing unit. The specification unit specifies a hair style that is estimated to be suitable for a face of a user, by using a learning model that is a model having learned a relationship between information indicating the characteristics of a face and a hair style evaluated to suit the face. The acquisition unit obtains a posted image including a specified hair image that is an image of the hair style specified by the specification unit. The image processing unit generates a composite image that is obtained by combining the image of the hairstyle specified by the specification unit and the face image of the user, based on the posted image obtained by the acquisition unit and the face image of the user.SELECTED DRAWING: Figure 8
Owner:LY CORP

Matching audio fingerprints

Methods, apparatus, systems and articles of manufacture are disclosed to select reference sub-fingerprints for comparison to query sub-fingerprints based on a determination that a query sub-fingerprint is a match with a reference sub-fingerprint, generate a count vector that stores total counts of matches between the query sub-fingerprints and different subsets of the reference sub-fingerprints, each of the different subsets being aligned to the query sub-fingerprints at a different offset from a reference point, each of the different offsets being mapped by the count vector to a different total count, calculate a maximum count among the total counts, a median of the total counts, and a difference between the maximum count and the median of the total counts, and classify the reference sub-fingerprints as a match with the query sub-fingerprints based on the difference between the maximum count in the count vector and the median.
Owner:GRACENOTE INC

Matching video content with podcast episodes

A system and method are provided for matching videos and podcast episodes. A data store comprising podcast episode identifiers is accessed. The podcast episode identifiers are associated with one or more podcast episode attributes. A video content item is identified. The video content item includes one or more video content item attributes. A matching podcast episode identifier that matches the video content item is determined based on the one or more podcast episode attributes and the one or more video content item attributes. A ranking of one of the video content items or the matching podcast episode identifiers is adjusted to reflect a correspondence between the video content item and the matching podcast episode identifier. Information associated with the matching podcast episode identifier is provided to a first user device.
Owner:GOOGLE LLC

system

We provide the system. [Solution] Means for acquiring audio information, A means of converting this audio information into text information, A means of analyzing the converted text information to understand the content of the inquiry or potential threat, A means of generating appropriate responses or warnings based on inquiries, A means for converting the generated text response into audio information, Means for making these operations compatible with multiple languages, A system that includes means for collecting and continuously learning from evaluation results of responses and warnings.
Owner:SOFTBANK GROUP CORP

Audio playing method, device, medium and computing device

Embodiments of the present disclosure provide an audio playing method, device, medium and computing device, the method comprising: in response to a playing operation of a first audio, determining user information corresponding to a user triggering the playing operation; in response to determining that the user does not have complete audio playing permission of the first audio according to the user information, and the user satisfies complete audio audition condition of the first audio, playing the complete first audio. In the present disclosure, when the user satisfies the complete audio audition condition, the audio to which the playing operation of the user is directed is played completely, so that the user can listen to the played audio completely, thereby providing a more objective decision basis for the user to obtain complete audio playing permission, increasing user stickiness and improving user experience.
Owner:HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD

Search method, device, equipment and readable storage medium

The application discloses a retrieval method, device and equipment and a readable storage medium. The method comprises the following steps: obtaining first feature information of to-be-retrieved content, and performing retrieval in a first feature library to obtain a first result set; when entries in the first result set are first-type entries, obtaining second feature information of the entries; performing retrieval in a second feature library containing second-type entry features according to the second feature information to obtain a second result set; and outputting a matched result sequence according to the first result set and the second result set. The application can output a matched list containing specific entries such as original content, improve the retrieval experience of users, and ensure the exposure rate of original content in a platform.
Owner:HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD

System for the joint playback of playlists for vehicles and a method for its application

This document describes a system (100) for the collaborative playback of media playlists in vehicles. The system (100) comprises a server (102) that communicates with a vehicle infotainment unit (106). The infotainment unit (106) connects mobile devices (108), assigned to users (110), to the server (102) after user (110) authentication. The server (102) allows users (110) to select infotainment media to be played on the infotainment unit (106) using the mobile devices (108) and to create an initial collaborative playlist. This playlist contains the selected infotainment media, which are queued in a predefined configuration, as well as details about the respective users (110) who selected the infotainment media.The server (102) also plays the infotainment media associated with the created shared playlist when one of the users (110) has given their consent, and furthermore restricts the modification of the created shared playlist once the first created shared playlist has been selected and approved by the relevant users (110).
Owner:MERCEDES BENZ GROUP AG

Client-calculated content metadata-based voice inquiry service quality (QoS)

To deal with a traffic abrupt increase in a processing stack of a server base.SOLUTION: A method includes receiving an automated speech recognition (ASR) request from a user device that includes a speech input captured by the user device and content metadata associated with the speech input. The content metadata is generated by the user device. The method also includes determining a priority score for the ASR request based on the content metadata. The method includes caching the ASR request in a pre-processing backlog of pending ASR requests each having a corresponding priority score. The pending ASR requests in the pre-processing backlog are ranked in order of the priority scores. The method also includes providing, from the pre-processing backlog, one or more of the pending ASR requests to a backend-side ASR module. Pending ASR requests associated with higher priority scores are processed before pending ASR requests associated with lower priority scores.SELECTED DRAWING: Figure 1
Owner:GOOGLE LLC

Multimodal semantic analysis and image search

A system and method are provided for identifying and retrieving semantically similar images from a database. A visual language model is used to perform semantic analysis of the input query (406) and identify semantic concepts related to the input query. For the identified semantic concepts, a preliminary set of images is retrieved from the database (804). Related concepts are extracted from the images by comparing the images to a predefined label space using a tokenizer and identifying related concepts (806). A ranked list of related concepts is generated based on their frequency of occurrence in the set (810). By combining the input query with a specific related concept, a specific related concept is selected from the ranked list, and the preliminary set of images is narrowed down (812). Further semantic analysis is performed iteratively until a threshold condition is met (814), and an additional set of images semantically similar to the combined input query and the selection of a specific related concept is retrieved.
Owner:NEC LABORATORIES AMERICA INC

English pronunciation calibration device and system

The present invention provides an English pronunciation calibration device, which includes: a memory, a word playback component, an image acquisition component, and a pronunciation correction processor. The memory stores a pre-stored word library. The word playback component has a display screen connected to the memory. The word playback component obtains the current word from the pre-stored word library and displays the current word on the display screen. The image acquisition component includes: a fixed plate, two fixed clamping rods, a camera, an audio acquisition module, and a wireless transmission module. By monitoring the learner's mouth shape during pronunciation, pronunciation defects are corrected. While providing the learner with English text and standard English pronunciation, the learner is also provided with English mouth shape differences corresponding to the standard pronunciation. By enabling the learner to obtain more relevant information, the learner's learning efficiency and English pronunciation accuracy are improved. The present invention also provides an English pronunciation calibration system.
Owner:JIANGXI NORMAL UNIV

Method, system, and computer-readable medium for comparing audio files and audio samples

The present invention relates to a method for comparing an audio file and an audio sample, comprising: S101: obtaining a complex frequency spectrum of the audio file; S102: obtaining a self-coherent sequence of the audio file and a deformed audio, wherein the deformed audio is obtained based on the audio file; S103: obtaining a coherent time series between the audio sample and the audio file; S104: performing deconvolution processing on the coherent time series using the self-coherent time series as a deconvolution kernel; S105: locating the audio file and / or the audio sample based on the deconvolved coherent time series. In the above embodiment of the present invention, the coherent time series of the audio sample and the audio file is deconvolved using the self-coherent time series of the audio file, which can more accurately locate the time position of the retrieved audio. After actual verification, the embodiment of the present invention has been verified to have good robustness in actual complex scenarios (such as low signal-to-noise ratio environments).
Owner:CHONGJI TECH BEIJING CO LTD

system

Provide a system. 【Solution means】 Means for receiving voice information and obtaining data, Voice recognition means for converting the data into character information, Analysis means for analyzing the converted character information and extracting a procedure flow, Generating means for automatically generating a procedure flow diagram based on the analysis result, Means for presenting operational issues and countermeasures based on the procedure flow diagram, Means for automatically generating materials based on the procedure flow diagram and countermeasures, Robot control means for updating an operation procedure through the generating means, A system including the above.
Owner:SOFTBANK GROUP CORP

system

The system according to this embodiment aims to generate and provide music to the user based on the atmosphere and emotions of a photograph. [Solution] The system according to the embodiment comprises a reception unit, an analysis unit, a generation unit, and a provision unit. The reception unit receives photos uploaded by the user. The analysis unit detects the characteristics of the photos received by the reception unit and analyzes the atmosphere, color tone, and emotion. The generation unit generates music based on the results analyzed by the analysis unit. The provision unit provides the music generated by the generation unit.
Owner:SOFTBANK GROUP CORP

Audio processing method and apparatus

An audio processing method and an electronic apparatus are provided. The audio processing method is applied to a conference system, and the conference system includes at least one audio capturing device. The audio processing method includes: receiving at least one segment of audio captured by the at least one audio capturing device; determining voices of a plurality of targets in the at least one segment of audio; and performing voice recognition on a voice of each of the plurality of targets, to obtain semantics corresponding to the voice of each target. Voice recognition is separately performed on voices of different targets, thereby improving accuracy of voice recognition.
Owner:HUAWEI TECH CO LTD +1

Information processing device, information processing method, and program

Provided are a device and method for executing a search process for data of a category different from text, using a query and word-level weight information. This information processing device includes a data processing unit into which a query composed of text applied to a data search process is input, and which executes a data search process based on the query. Together with the query, word-level weight information constituting the query are input into the data processing unit, which executes a search process for data of a category different from the text using the query and the word-level weight information. For a word assigned with a positive weight, a search process with increased influence is executed, and for a word assigned with negative weight, a search process with reduced influence is executed.
Owner:SONY GROUP CORP

Internet news information capturing and audio abstract broadcasting method

The invention relates to the technical field of information capturing, in particular to an internet news information capturing and audio abstract broadcasting method. A cloud server firstly captures news data from each Internet information source through a web crawler tool, generates a corresponding news abstract audio, and obtains a corresponding content tag; the cloud server calculates a preference matching value of the current user for each content tag based on listening data of the current user in past first preset duration, and the higher the preference matching value is, the higher the preference degree of the current user for the news abstract audio of the corresponding content tag is; then, based on the content labels of the news summary audios and the preference matching values of the current user for all the content labels, a listening sequence of the current user for listening to the news summary audios is determined, and a subsequent user terminal installs the listening sequence to obtain and play all the news summary audios from the cloud server; therefore, the news audio which the user is interested in can be accurately identified and pushed.
Owner:HUNAN MANGO INTELLIGENT MEDIA TECH DEV CO LTD