Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

7 results about "Speech retrieval" patented technology

Video content retrieval method, device and terminal based on voice interaction of television system

The invention discloses a video content retrieval method and device based on television system voice interaction and a terminal, and relates to the technical field of video processing, and the method comprises the steps: when a video is played for the first time, extracting a picture frame from the video at a preset frequency, converting the picture frame into a multi-dimensional image feature vector and a corresponding video timestamp, and carrying out hierarchical storage in a database, constructing vectorized data containing visual semantic information; obtaining a voice retrieval instruction, performing intention recognition and semantic understanding, extracting a detection keyword, and generating a multi-dimensional retrieval feature vector; calculating a matching degree between the multi-dimensional retrieval feature vector and a multi-dimensional image feature vector of a video picture frame stored in a database, and screening out picture frames of which the similarity is higher than a preset similarity threshold to form a retrieval candidate matching set; and determining matched picture playing. The video content retrieval method is efficient, accurate and high in interactivity, and retrieval experience and operation efficiency of the user in the video watching process are remarkably improved.
Owner:SHENZHEN COOCAA NETWORK TECH CO LTD

Voice retrieval device and method based on intelligent AI technology

The invention discloses voice retrieval equipment and method based on an intelligent AI technology, and relates to the technical field of voice retrieval, the voice retrieval equipment comprises a shell, a top groove is formed in the top of the shell, a moving groove is formed in the shell, and a first servo motor is fixedly connected to one side of an inner cavity of the moving groove; according to the voice retrieval equipment based on the intelligent AI technology and the voice retrieval method based on the intelligent AI technology, by arranging the shell, the bottom plate and the suction cup, the equipment can be limited at the wrist of a user to be used along with the user; in addition, the equipment can be installed on the desktop of the user for use, and when the equipment is installed on the desktop of the user for use, the shell is adsorbed on the desktop of the user, so that the shell is prevented from moving.
Owner:SHENZHEN ENERGY BRIGHT POWER CO LTD

Video clip retrieval processing method based on voice interaction and terminal

The invention discloses a video clip retrieval processing method based on voice interaction and a terminal, and relates to the technical field of intelligent terminals, and the method comprises the following steps: obtaining a voice retrieval instruction for retrieving a video clip; analyzing the voice retrieval instruction to obtain a video clip retrieval requirement; generating a video clip structured retrieval condition according to the analyzed video clip retrieval requirement; frame extraction analysis is carried out on a video needing to be retrieved, and a correlation video clip group is constructed based on continuous frame correlation and a preset duration constraint condition; according to the generated video clip structured retrieval condition, retrieving a retrieval clip corresponding to the video clip retrieval requirement from the constructed correlation video clip group; and outputting the retrieved retrieval fragment corresponding to the video fragment retrieval requirement. According to the method and the device, fragment-level retrieval is realized, continuous fragments with highly related contents can be accurately identified, the fragment range is displayed through differentiated UI identifiers, and the retrieval efficiency and the interaction experience of a user are remarkably improved.
Owner:SHENZHEN COOCAA NETWORK TECH CO LTD

System and method for assessing and correcting potential underserved content in natural language understanding applications

Methods, systems, and related products that provide detection of media content items that are under-locatable by machine voice-driven retrieval of uttered requests for retrieval of the media items. For a given media item, a resolvability value and / or an utterance resolve frequency is calculated by a number of playbacks of the media item by a speech retrieval modality to a total number of playbacks of the media item regardless of retrieval modality. In some examples, the methods, systems and related products also provide for improvement in the locatability of an under-locatable media item by collecting and / or generating one or more pronunciation aliases for the under-locatable item.
Owner:SPOTIFY

Speech retrieval method, system and medium based on semantic analysis and high-dimensional modeling

The application discloses a speech retrieval method and system based on semantic analysis and high-dimensional modeling and a medium, and the method comprises the following steps: acquiring speech text data; synchronizing the text data for constructing semantic labels and defining data models to a retrieval database through data synchronization; determining search text and screening condition text data for high screening processing; and performing text recall, and generating a final recall candidate set after sorting processing and filtering processing. The application can comprehensively integrate the respective characteristics of three recall modes, and improve the accuracy of text recall as much as possible. The application introduces retrieval based on multi-dimensional semantic labels and text length, realizes recall based on context information, solves the problem that context information cannot be considered in a general text recall method, and improves retrieval efficiency while ensuring retrieval accuracy.
Owner:GUANGZHOU TANJI TECH CO LTD

A method for enhancing robustness of a voiceprint retrieval model

ActiveCN116312550BEngineeringData mining
This invention discloses a method for enhancing the robustness of a voiceprint retrieval model, belonging to the field of voiceprint retrieval. The invention aims to provide a method to enhance the robustness of systems using deep learning to build speech retrieval models, helping the system maintain relatively good performance even when subjected to adversarial example attacks. First, during a retrieval process, the relevance of the search results is ranked, and relatively relevant speech is selected. Then, the query statement is denoised. Next, using the denoised query speech, a second query is performed among the previously selected relatively relevant speech, selecting the most relevant speech from this second query as the final query result. This invention provides a method to enhance the robustness of voiceprint retrieval models in deep learning-based voiceprint retrieval systems, helping the system operate more effectively.
Owner:NANJING UNIV

Nuclear emergency medical rescue intelligent voice auxiliary system

The invention provides an intelligent voice auxiliary system for nuclear emergency medical rescue, and relates to the technical field of natural voice processing. The system comprises a voice recognition module used for performing role separation and fragment recombination on sound wave signals in a nuclear emergency medical teaching scene to obtain a plurality of voice signals; wherein each voice signal corresponds to one speaking role; the voice retrieval module is used for identifying terminologies in each voice signal based on a system knowledge resource library and a practical training scene library; the voice evaluation module is used for evaluating the instruction specification degree of each voice signal based on a preset nuclear medicine emergency rescue knowledge graph and the terminology in each voice signal; wherein the nuclear medicine emergency rescue knowledge graph comprises a role entity, a process entity, an injury condition entity and a material entity. According to the invention, the problem of multi-role voice confusion when a virtual system is used for nuclear emergency medical teaching and training can be solved, the accuracy of professional term recognition is improved, and efficient and accurate voice interaction support is provided for theoretical teaching and practical training.
Owner:THE FIFTH MEDICAL CENT OF CHINESE PLA GENERAL HOSPITAL