Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

18 results about "Voice Tag" patented technology

Voice tags are used in automated speech recognition in a voice command device, allowing the user to "speak" commands. For example, using voice commands with an automated device, such as an IVR telephone prompt or to dial a contact on a mobile phone.

Voice tag generation method and device, electronic equipment and storage medium

The invention relates to a voice tag generation method and device, electronic equipment and a storage medium, and belongs to the technical field of computers, and the method comprises the steps: carrying out the feature matrix decomposition of a plurality of to-be-processed first voice feature matrixes, and obtaining a decomposed first matrix and a decomposed second matrix; performing hierarchical clustering on column vectors in the first matrix to obtain a tag tree of the first matrix; performing projection processing on the second voice according to the second matrix to obtain a voice feature representation vector of which the dimension is the same as the column number of the first matrix; querying a tag tree, and determining a feature tag of the second voice; wherein the feature tag of the second voice is determined according to the clustering tag in the tag tree mapped by the voice feature representation vector. According to the embodiment of the invention, the construction and maintenance cost of a voice feature tag generation system can be remarkably reduced, and the automation degree and robustness of tag generation are improved.
Owner:MOORE THREADS TECH CO LTD

Method and device for documenting care measures

A method for the automated documentation of care measures for the care of a patient has the steps of: capturing image data of interactions between a caregiver and a patient using an image capture unit, processing the captured image data using an image processing algorithm for recognizing the interactions in the image data, and generating text and / or speech labels for the recognized interactions using a machine label generation learning model for the automated documentation of care measures using the generated text and / or speech labels.
Owner:FORBENCAP GMBH

Text history-based reasoning method and device in multi-modal dialogue scene, medium and program product

The invention belongs to the technical field of natural language processing, and provides a text history-based reasoning method and device in a multi-modal dialogue scene, a medium and a program product, and the method comprises the steps: obtaining a historical dialogue of voice and text alternation, and carrying out the automatic voice recognition of a voice part in the historical dialogue to obtain a recognition text; extracting non-text feature information at least containing mood, emotion or emphasis when the recognition text is generated, compressing and encoding the non-text feature information into a voice tag, and adding the voice tag to the corresponding recognition text to form historical text information with a semantic compensation mark; and then combining the historical text information and the current user input voice into model input content, inputting the model input content into a multi-modal dialogue model for reasoning, and outputting a corresponding text reply. According to the method, the length of a model input sequence can be shortened, the reasoning efficiency is improved, meanwhile, key non-text information is reserved, the dialogue accuracy is improved, and the multi-modal dialogue interaction experience is optimized.
Owner:BEIJING BAIGEFEICHI TECH LLC

Conversation system for voice control of motor vehicle

The invention relates to a dialogue system (100) for voice control of a motor vehicle (105), the dialogue system (100) comprising a device (110) for detecting a voice input (315, 325, 335, 345) of a person (310, 320, 330, 340) on the motor vehicle (105); and a processing device (120) arranged to recognize a language of the speech input (315, 325, 335, 345) on the basis of the speech marker and to provide a response (350) to the speech input (315, 325, 335, 345), the response (350) being provided in the recognized language.
Owner:BAYERISCHE MOTOREN WERKE AG

Effective speech recognition method and device

The application provides an effective speech recognition method and device, the method comprises the following steps: based on an effective speech recognition model, extracting audio features of to-be-recognized audio data, and determining effective speech data from the to-be-recognized audio data by applying the audio features of the to-be-recognized audio data; the effective speech recognition model takes minimizing the difference between effective predicted speech and effective speech labels, minimizing the distance between the audio features of sample audio data and the audio features of the sample audio data after adding noise, and maximizing the distance between the audio features of the sample audio data and the audio features of pure noise data as a training target, and the effective predicted speech is obtained by the effective speech recognition model on the sample audio data. In the face of a scene with small speech signal-to-noise ratio and large background noise, the application can accurately recognize the effective speech of the to-be-recognized audio data, and improve the effective speech recognition precision.
Owner:HEFEI IFLY DIGITAL TECH CO LTD

A voice activity detection method, system, terminal and storage medium

The application provides a voice activity detection method, system, terminal and storage medium, relates to a voice data processing technical field, and in particular to a voice activity detection method, which comprises the following steps: obtaining a voice training sample carrying a voice label; performing feature extraction according to the voice training sample to obtain a root amplitude spectrum feature sample and an Fbank feature sample; and constructing a voice activity detection model, wherein the voice activity detection model is used for performing feature fusion on the root amplitude spectrum feature and the Fbank feature obtained by performing feature extraction on voice data to obtain fusion features, and outputting a probability value of each frame of the voice data existing a voice signal based on the fusion features; training the voice activity detection model by using the root amplitude spectrum feature sample and the Fbank feature sample to obtain a trained voice activity detection model; and performing voice activity detection by using the trained voice activity detection model. The application can improve the accuracy of voice activity detection.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Neural network-based synthesized speech method, system, device, and storage medium

This invention relates to the field of artificial intelligence technology, and proposes a speech synthesis method, system, device, and storage medium based on neural networks. The method includes: acquiring target text information; inputting the target text information into a speech synthesis model for processing to obtain the target Mel spectrum corresponding to the target text information. The speech synthesis model is an autoregressive speech synthesis model, trained using training text and speech tags. During training, the speech synthesis model aligns the target text information and the target speech based on the linear relationship between the training text and the speech tags; and generates the target speech information based on the target Mel spectrum. This invention improves the training efficiency of the autoregressive speech synthesis model, mitigates word omission and repetition issues, and enhances the robustness of the speech synthesis method by introducing prior information about the linear relationship between the training text and the speech tags during the speech synthesis model training process.
Owner:PING AN TECH (SHENZHEN) CO LTD

Model training method and device, speech synthesis method, equipment and storage medium

Embodiments of the present application provide a model training method and device, a speech synthesis method, equipment and a storage medium, and relate to the technical field of artificial intelligence. The model training method obtains a training data set, obtains a monotonic alignment loss function used for training an attention unit and a preset loss function used for training a speech synthesis model, trains the speech synthesis model based on the preset loss function in combination with a speech output vector and a corresponding speech label, adjusts model parameters of the speech synthesis model, and obtains a trained speech synthesis model when a value of the loss function meets a preset condition. In the training process, the monotonicity of an attention weight sequence is trained based on the monotonic alignment loss function. The monotonic alignment loss function is set for attention alignment, the monotonicity of the attention weight sequence is ensured, which helps to realize fast convergence of the model, improve the training accuracy of the model, and improve the naturalness and robustness of the synthesized language.
Owner:PING AN TECH (SHENZHEN) CO LTD

Neurally guided speech enhancement method based on personalized brain electrode distribution

The present application provides a neural-guided speech enhancement method based on personalized brain electrode distribution. In the speech enhancement method, a training method of a speech enhancement model includes obtaining a training set, the training set including multiple training sample combinations corresponding to at least one training user, the training sample combinations including the training user's real biological auxiliary information, mixed distribution auxiliary information, Gaussian distribution auxiliary information, mixed training speech information and enhanced speech labels, the auxiliary information representing the degree of attention of the training user to a certain identified user in the mixed training speech information; using the real biological auxiliary information, mixed distribution auxiliary information, Gaussian distribution auxiliary information and mixed training speech information to perform adversarial training on multiple initial enhancement models to obtain multiple enhanced speech information; generating an initial loss value based on the enhanced speech information and the enhanced language label; iteratively adjusting the network parameters of the initial enhancement model based on the initial loss value to obtain a speech enhancement model.
Owner:UNIV OF SCI & TECH OF CHINA

Speech recognition method, speech recognition model training method, device, and medium

This application discloses a speech recognition method, a speech recognition model training method, an apparatus, and a medium. The method includes acquiring a speech to be recognized and acquiring a trained speech recognition model, the speech recognition model including an encoding network and a decoding network. At each stage of encoding the speech to be recognized using the encoding network, the speech to be recognized is first classified by a target speech attribute to obtain a predicted attribute category to which the speech to be recognized belongs, then encoded based on the predicted attribute category of the target speech attribute, obtaining a first encoding feature, and decoding the first encoding feature based on the decoding network to obtain a recognized text of the speech to be recognized. The speech recognition model is adjusted based on at least a first loss, the first loss characterizing the difference between a preset attribute category of a sample speech label with the target speech attribute and the attribute category of a sample recognized by the speech recognition model. In this way, the application can improve speech recognition accuracy while reducing costs.
Owner:IFLYTEK CO LTD

An AI intelligent calling method

The present invention provides an AI intelligent calling method, comprising the steps of: obtaining voice content when promoting a current product to a current user, obtaining a call time period, purchase intention, and the current user's call scenario based on the voice content; obtaining a user profile of the current user, correlating the above analysis content, and obtaining a voice tag; when it is necessary to promote a target product, matching a first target tag from all voice tags, determining whether the number of current users in the first target tag reaches a target number, and if so, using the current user in the first target tag as the target user; otherwise, matching a second target tag based on the user profile and the call scenario, and combining the two as the target user; calling the target user based on the call time period in the target tag corresponding to the target user to promote the target product. The present invention can ensure the accuracy and communication rate of voice calls.
Owner:FUJIAN BOSHITONG INFORMATION CO LTD

Vehicle-mounted multi-zone speech separation method, electronic device, and storage medium

The present invention discloses a vehicle-mounted multi-zone speech separation method, an electronic device and a storage medium, wherein a vehicle-mounted multi-zone speech separation method includes: convolving acquired high-fidelity audio with acquired room impulse response data to obtain a mixed signal and at least one speech label; training a fusion beamforming network model based on the mixed signal and the at least one speech label; testing the fusion beamforming network model based on a preset simulation test set to determine whether the fusion beamforming network model meets preset requirements; if the preset requirements are met, predicting the beamforming weights of the mixed signal and the at least one speech label based on the fusion beamforming network model to obtain each sound zone separation signal.
Owner:AISPEECH CO LTD

Method and apparatus for documentation of an operation on a patient

A method for automated documentation of an operation of a patient which may have the steps of: capturing image data of interactions between a surgeon and / or a surgical utensil and a patient using an image capturing unit; processing the captured image data using an image processing algorithm to recognize the interactions in the image data; and generating text and / or speech labels for the recognized interactions using a machine learning label generation model for automated documentation of the surgery using the generated text and / or speech labels.
Owner:FORBENCAP GMBH

Method and device for documenting an operation on a patient

Method for the automated documentation of a patient's operation (120), comprising the steps: capturing (S1) image data of interactions (130) between a surgeon and / or a surgical instrument (110) and a patient (120) using an image acquisition unit (10); processing (S2) the captured image data using an image processing algorithm to detect the interactions in the image data; and generating (S3) text and / or speech labels for the detected interactions (130) using a machine label generation learning model (31) for automated documentation of the operation using the generated text and / or speech labels.
Owner:FORBENCAP GMBH

Method for training a speech recognition model and speech recognition method

ActiveCN115691475BBiological modelsSpeech recognitionSpeech trainingSpeech sound
The application relates to a method for training a speech recognition model, comprising: providing a speech training data set comprising a plurality of speech data and speech labels corresponding to each speech data; providing a speech recognition model to be trained, the speech recognition model comprising a convolutional neural network, a first fully connected network, a recurrent neural network and a second fully connected network coupled in cascade, wherein each network comprises one or more network layers with a parameter matrix; wherein the speech recognition model is used to process speech data to generate corresponding speech recognition results; and training the speech recognition model using the speech training data set, so that after training, the parameter matrices of at least two adjacent network layers in the speech recognition model satisfy a predetermined constraint condition; and so that the accuracy of the speech recognition results of the speech recognition model on the speech data calculated using at least one loss function satisfies a predetermined recognition target.
Owner:MONTAGE TECH CHENGDU CO LTD

Procedure and device for documenting nursing measures

Method for the automated documentation of nursing measures for the care of a patient, comprising the steps: capturing (S1) image data of interactions between a nurse and a patient using an image acquisition unit, processing (S2) the captured image data using an image processing algorithm to recognize the interactions in the image data, and generating (S3) text and / or speech labels for the recognized interactions using a machine label generation learning model for the automated documentation of nursing measures using the generated text and / or speech labels.
Owner:FORBENCAP GMBH

Method and system for displaying running state data of water turbine equipment of hydropower station by using voice instruction

The invention belongs to the technical field of hydropower station management, and particularly provides a method and system for displaying hydropower station water turbine equipment operation state data through voice instructions, and the method comprises the steps: obtaining real-time equipment data and three-dimensional map display data; the method comprises the following steps: sampling a voice signal according to a set sampling rate, converting the voice signal into time sequence waveform data, and then performing de-noising processing on the voice signal; a CNN-LSTM model is adopted to convert voice input of a user into a primary text; constructing a voice tag library and a vocabulary library, and performing instruction analysis on the text of the output voice recognition result by adopting a BERT pre-training model; and according to the instruction analysis result, obtaining the operation state data of the corresponding equipment to display the operation state of the equipment, and / or obtaining the three-dimensional map data to display a three-dimensional map. According to the method and the system, the hydroelectric equipment operation state is displayed according to the user voice instruction, and the intelligent management level is improved.
Owner:CHINA YANGTZE POWER

Speech keyword detection method and apparatus, device, and storage medium

The embodiment of the application provides a voice keyword detection method and device, electronic equipment and a storage medium, belongs to the technical field of voice processing, and is suitable for the field of financial technology. The method comprises the following steps: acquiring voice sample data; acquiring an initial feature extraction submodel and an initial adversarial classification submodel; performing keyword detection on the voice sample data based on the initial feature extraction submodel to obtain a predicted keyword; calculating feature extraction loss data based on the predicted keyword and a key frame label; performing type identification on the voice sample data based on the initial adversarial classification submodel and sample voice features to obtain a voice prediction category; calculating adversarial classification loss data based on the voice prediction category and a voice label; and performing gradient inversion on the initial feature extraction submodel and the initial adversarial classification submodel based on the feature extraction loss data and the adversarial classification loss data. The embodiment of the application can reduce the cost consumption of voice keyword detection.
Owner:PING AN TECH (SHENZHEN) CO LTD