Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

11 results about "Voice Tag" patented technology

Voice tags are used in automated speech recognition in a voice command device, allowing the user to "speak" commands. For example, using voice commands with an automated device, such as an IVR telephone prompt or to dial a contact on a mobile phone.

Voice tag generation method and device, electronic equipment and storage medium

The invention relates to a voice tag generation method and device, electronic equipment and a storage medium, and belongs to the technical field of computers, and the method comprises the steps: carrying out the feature matrix decomposition of a plurality of to-be-processed first voice feature matrixes, and obtaining a decomposed first matrix and a decomposed second matrix; performing hierarchical clustering on column vectors in the first matrix to obtain a tag tree of the first matrix; performing projection processing on the second voice according to the second matrix to obtain a voice feature representation vector of which the dimension is the same as the column number of the first matrix; querying a tag tree, and determining a feature tag of the second voice; wherein the feature tag of the second voice is determined according to the clustering tag in the tag tree mapped by the voice feature representation vector. According to the embodiment of the invention, the construction and maintenance cost of a voice feature tag generation system can be remarkably reduced, and the automation degree and robustness of tag generation are improved.
Owner:MOORE THREADS TECH CO LTD

Method and device for documenting care measures

A method for the automated documentation of care measures for the care of a patient has the steps of: capturing image data of interactions between a caregiver and a patient using an image capture unit, processing the captured image data using an image processing algorithm for recognizing the interactions in the image data, and generating text and / or speech labels for the recognized interactions using a machine label generation learning model for the automated documentation of care measures using the generated text and / or speech labels.
Owner:FORBENCAP GMBH

Text history-based reasoning method and device in multi-modal dialogue scene, medium and program product

The invention belongs to the technical field of natural language processing, and provides a text history-based reasoning method and device in a multi-modal dialogue scene, a medium and a program product, and the method comprises the steps: obtaining a historical dialogue of voice and text alternation, and carrying out the automatic voice recognition of a voice part in the historical dialogue to obtain a recognition text; extracting non-text feature information at least containing mood, emotion or emphasis when the recognition text is generated, compressing and encoding the non-text feature information into a voice tag, and adding the voice tag to the corresponding recognition text to form historical text information with a semantic compensation mark; and then combining the historical text information and the current user input voice into model input content, inputting the model input content into a multi-modal dialogue model for reasoning, and outputting a corresponding text reply. According to the method, the length of a model input sequence can be shortened, the reasoning efficiency is improved, meanwhile, key non-text information is reserved, the dialogue accuracy is improved, and the multi-modal dialogue interaction experience is optimized.
Owner:BEIJING BAIGEFEICHI TECH LLC

Conversation system for voice control of motor vehicle

The invention relates to a dialogue system (100) for voice control of a motor vehicle (105), the dialogue system (100) comprising a device (110) for detecting a voice input (315, 325, 335, 345) of a person (310, 320, 330, 340) on the motor vehicle (105); and a processing device (120) arranged to recognize a language of the speech input (315, 325, 335, 345) on the basis of the speech marker and to provide a response (350) to the speech input (315, 325, 335, 345), the response (350) being provided in the recognized language.
Owner:BAYERISCHE MOTOREN WERKE AG

A voice activity detection method, system, terminal and storage medium

The application provides a voice activity detection method, system, terminal and storage medium, relates to a voice data processing technical field, and in particular to a voice activity detection method, which comprises the following steps: obtaining a voice training sample carrying a voice label; performing feature extraction according to the voice training sample to obtain a root amplitude spectrum feature sample and an Fbank feature sample; and constructing a voice activity detection model, wherein the voice activity detection model is used for performing feature fusion on the root amplitude spectrum feature and the Fbank feature obtained by performing feature extraction on voice data to obtain fusion features, and outputting a probability value of each frame of the voice data existing a voice signal based on the fusion features; training the voice activity detection model by using the root amplitude spectrum feature sample and the Fbank feature sample to obtain a trained voice activity detection model; and performing voice activity detection by using the trained voice activity detection model. The application can improve the accuracy of voice activity detection.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Neural network-based synthesized speech method, system, device, and storage medium

This invention relates to the field of artificial intelligence technology, and proposes a speech synthesis method, system, device, and storage medium based on neural networks. The method includes: acquiring target text information; inputting the target text information into a speech synthesis model for processing to obtain the target Mel spectrum corresponding to the target text information. The speech synthesis model is an autoregressive speech synthesis model, trained using training text and speech tags. During training, the speech synthesis model aligns the target text information and the target speech based on the linear relationship between the training text and the speech tags; and generates the target speech information based on the target Mel spectrum. This invention improves the training efficiency of the autoregressive speech synthesis model, mitigates word omission and repetition issues, and enhances the robustness of the speech synthesis method by introducing prior information about the linear relationship between the training text and the speech tags during the speech synthesis model training process.
Owner:PING AN TECH (SHENZHEN) CO LTD

Model training method and device, speech synthesis method, equipment and storage medium

Embodiments of the present application provide a model training method and device, a speech synthesis method, equipment and a storage medium, and relate to the technical field of artificial intelligence. The model training method obtains a training data set, obtains a monotonic alignment loss function used for training an attention unit and a preset loss function used for training a speech synthesis model, trains the speech synthesis model based on the preset loss function in combination with a speech output vector and a corresponding speech label, adjusts model parameters of the speech synthesis model, and obtains a trained speech synthesis model when a value of the loss function meets a preset condition. In the training process, the monotonicity of an attention weight sequence is trained based on the monotonic alignment loss function. The monotonic alignment loss function is set for attention alignment, the monotonicity of the attention weight sequence is ensured, which helps to realize fast convergence of the model, improve the training accuracy of the model, and improve the naturalness and robustness of the synthesized language.
Owner:PING AN TECH (SHENZHEN) CO LTD

Speech recognition method, speech recognition model training method, device, and medium

This application discloses a speech recognition method, a speech recognition model training method, an apparatus, and a medium. The method includes acquiring a speech to be recognized and acquiring a trained speech recognition model, the speech recognition model including an encoding network and a decoding network. At each stage of encoding the speech to be recognized using the encoding network, the speech to be recognized is first classified by a target speech attribute to obtain a predicted attribute category to which the speech to be recognized belongs, then encoded based on the predicted attribute category of the target speech attribute, obtaining a first encoding feature, and decoding the first encoding feature based on the decoding network to obtain a recognized text of the speech to be recognized. The speech recognition model is adjusted based on at least a first loss, the first loss characterizing the difference between a preset attribute category of a sample speech label with the target speech attribute and the attribute category of a sample recognized by the speech recognition model. In this way, the application can improve speech recognition accuracy while reducing costs.
Owner:IFLYTEK CO LTD

Method and apparatus for documentation of an operation on a patient

A method for automated documentation of an operation of a patient which may have the steps of: capturing image data of interactions between a surgeon and / or a surgical utensil and a patient using an image capturing unit; processing the captured image data using an image processing algorithm to recognize the interactions in the image data; and generating text and / or speech labels for the recognized interactions using a machine learning label generation model for automated documentation of the surgery using the generated text and / or speech labels.
Owner:FORBENCAP GMBH

Method and device for documenting an operation on a patient

Method for the automated documentation of a patient's operation (120), comprising the steps: capturing (S1) image data of interactions (130) between a surgeon and / or a surgical instrument (110) and a patient (120) using an image acquisition unit (10); processing (S2) the captured image data using an image processing algorithm to detect the interactions in the image data; and generating (S3) text and / or speech labels for the detected interactions (130) using a machine label generation learning model (31) for automated documentation of the operation using the generated text and / or speech labels.
Owner:FORBENCAP GMBH

Procedure and device for documenting nursing measures

Method for the automated documentation of nursing measures for the care of a patient, comprising the steps: capturing (S1) image data of interactions between a nurse and a patient using an image acquisition unit, processing (S2) the captured image data using an image processing algorithm to recognize the interactions in the image data, and generating (S3) text and / or speech labels for the recognized interactions using a machine label generation learning model for the automated documentation of nursing measures using the generated text and / or speech labels.
Owner:FORBENCAP GMBH