Voice data acquisition method and system

A voice data and voice technology, applied in the field of voice data acquisition, can solve problems such as no solutions have been proposed, achieve the effect of enhancing specificity and security, reducing training burden, and wide application range

CN108510981AActive Publication Date: 2018-09-07SAMSUNG ELECTRONICS CHINA R&D CENT +1
6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Current Assignee / Owner
Publication Date
2018-09-07

Smart Images

  • Figure 1
    Figure 1
  • Figure 2
    Figure 2
  • Figure 3
    Figure 3
Patent Text Reader

Abstract

The invention provides a voice data acquisition method and a voice data acquisition system. The method comprises the following steps: in a process of voice communication of a user, saving a voice dataflow which is transmitted in an intelligent terminal system in real time, saving a voice data flow which is inputted from a microphone as first voice data and saving a voice data flow which is outputted from a headset as second voice data; detecting whether the first voice data and the second voice data conform to training requirements of a voice recognition model or not, if so, continuing to judge whether the first voice data is derived from an application object of the voice recognition model or not, if so, marking the first voice data as application object voice data and marking the secondvoice data as non-application object voice data; and if not, marking the first voice data and the second voice data as non-application object voice data. On the basis of the method provided by the invention, by the voice improving acquisition method, burden of a user in training the voice recognition model is reduced and user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

technical field

[0001] The invention relates to the field of artificial intelligence, in particular to a method and system for acquiring voice data. Background technique

[0002] Mobile terminal speech recognition is divided into two categories: semantic recognition and speaker recognition

[0003] Speaker recognition is often referred to as voiceprint recognition. It is generally divided into two types: text-dependent and text-independent.

[0004] Text-related speech recognition usually requires the user to repeat the fixed words and sentences 2-3 times. To record the relevant characteristic information as registration (Enroll). When in use, the user is also required to read the same fixed words and phrases for speech discrimination (Predict).

[0005] Non-text-related speech recognition does not require the user to follow a fixed sentence. Users input a large amount of voice data as machine learning training (Train), and the user's feature information is highly purif...

Examples

Embodiment Construction

[0029] In order to make the object, technical solution and advantages of the present invention clearer, the present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0030] figure 1 It is a flow chart of the method for obtaining voice data of the present invention, comprising the following steps:

[0031] Step A-1 (S101): When the user makes a voice call, save the voice data stream transmitted in real time in the smart terminal system, save the input voice data stream of the microphone as the first voice data, and save the output voice data stream of the handset as second voice data;

[0032] Step A-2 (S102): Detect whether the first speech data and the second speech data meet the speech recognition model training requirements, if so, perform step A-3;

[0033] Step A-3 (S103): judging whether the first speech data is from the application object of the speech recognition model, if yes, execute step A-4, if n...